Data Architect
Persistent Systems · Security DU8 · Lead
Apply ↗Skills
["Java""Security"]Technology stack
{}Description
About Persistent We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what’s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem. Our disruptor’s mindset, commitment to client success, and agility to thrive in the dynamic environment have enabled us to sustain our growth momentum. Persistent has been recognized across top industry platforms for innovation, leadership, and inclusion. We reported $1,654.4M FY26 revenue with 17.4% Y-o-Y growth. We have delivered 24 sequential quarters of growth with $436.0M in Q4 FY26 revenue, up 3.2% Q-o-Q and 16.2% Y-o-Y growth. Our 27,500+ global team members, located in 18 countries, have been instrumental in helping the market leaders transform their industries. We have been recognized as the Fastest Growing IT Services Brand Globally in the 2026 Brand Finance IT Services 25 Report. We named a Leader in the Everest Group Private Equity (PE) Services PEAK Matrix® Assessment 2026 and Software Product Engineering PEAK Matrix® Assessment 2026. About Position: We are seeking an experienced Lead Data Architect to drive a strategic IAM Data Modernization Program focused on migrating an on-premises SQL Data Warehouse into a modern, scalable Data Lake ecosystem on Google Cloud Platform (GCP). The ideal candidate will have strong expertise in Data Architecture, Big Data Engineering, Data Lake Design, Cloud Platforms, Metadata-Driven ETL Frameworks, PySpark, BigQuery, and Enterprise Data Modernization. This role will provide architectural leadership across OpenShift (OCP), S3, PostgreSQL metadata repositories, and GCP-based analytics platforms while enabling advanced reporting, analytics, and Generative AI use cases such as natural language querying, intelligent summarization, and cross-domain trend analysis. Role: Lead Data Architect Location: Bangalore Experience: 8 to 12 Years Job Type: Full-Time Employment What You'll Do: Own end-to-end architecture for the IAM Data Modernization Program. Define target-state architecture across OpenShift (OCP), S3 Data Lake, PostgreSQL Metadata Repository, GCP Data Platform, BigQuery Analytics Environment. Establish architecture standards, design guidelines, and governance frameworks. Drive enterprise data modernization and cloud migration initiatives. Architect and guide implementation of a reusable metadata-driven ETL framework. Define ingestion frameworks supporting Databases, Files, REST APIs, MongoDB, Splunk, Active Directory, Virtual Data Sources (VDS). Design metadata repositories using MD_* Static Metadata Tables, OPS_* Operational & Runtime Tables. Enable Pipeline Configuration, Run Tracking, Audit Logging, Checkpoint Recovery, Notifications, Schema Evolution Management. Design and implement scalable Data Lake architectures. Define storage layers including Landing Zone, Raw Zone, Archive Zone, Error Zone. Establish promotion, archival, recovery, and replay mechanisms. Define enterprise storage, retention, and lifecycle management strategies. Implement Data Lake best practices using layered architecture models such as Bronze, Silver, and Gold. Provide technical leadership for PySpark-based transformations and processing frameworks. Design scalable ETL/ELT pipelines for high-volume data ingestion and processing. Optimize distributed processing workloads for performance and scalability. Guide implementation of data quality and validation frameworks. Enable data processing for analytics, reporting, and AI workloads. Design future-state GCP analytics architecture. Define BigQuery dataset structures, schemas, partitioning, and clustering strategies. Architect secure integration between S3/OCP platforms and BigQuery. Develop performance optimization and cost management strategies. Establish cloud-native analytics and reporting capabilities. Define enterprise data governance standards and operational controls. Implement Data Quality Validation, Schema Validation, Schema Drift Detection, Checkpoint Restartability, Auditability, Operational Resilience, Data Lineage. Ensure compliance with security and governance requirements. Guide implementation of CI/CD automation using GitHub Actions, JFrog, Harness, Liquibase, Helm, OpenShift Deployments. Define deployment standards and environment promotion strategies. Establish DevOps governance and automation best practices. Enable enterprise analytics and GenAI capabilities through modern data architecture. Support Natural Language Querying, Advanced Trend Analysis, Intelligent Summarization, Cross-Domain Metrics Monitoring. Design data foundations for future AI, ML, and Generative AI solutions. Lead architecture discussions and design reviews. Provide technical guidance to Data Engineers and Development teams. Mentor teams on Python, PySpark, SQL, BigQuery, Data Modeling, Performance Optimization. Partner with client stakeholders to define technical roadmaps and implementation standards. Support architecture governance and production readiness reviews. Expertise You'll Bring: 8-12 years of experience in Data Architecture and Data Engineering. Strong expertise in PySpark, Apache Spark, ETL/ELT Frameworks, Distributed Data Processing, Data Transformation, Data Ingestion Frameworks. Experience building enterprise-scale data processing solutions. Strong experience designing cloud-based data platforms. Expertise in Google Cloud Platform (GCP), BigQuery, Cloud-Native Architectures, Data Warehousing, Enterprise Data Lakes. Experience leading large-scale data migration and modernization programs. Proven experience implementing Data Lake Architectures, Lakehouse Architectures, Bronze / Silver / Gold Models, Enterprise Storage Frameworks. Knowledge of GCS Design, Bucket Layout Strategy, Naming Standards, Lifecycle Policies, Retention Management. Database & Metadata Management using PostgreSQL SQL Development Data Modeling Metadata Repository Architecture Schema Design Performance Tuning Data Governance Experience integrating Relational Databases, APIs, Files, MongoDB, Splunk, Active Directory. Expertise in Full Load Processing, Incremental Loads, Watermark-Based Processing, CDC Patterns, Batch & API-Based Ingestion. GitHub Actions Harness JFrog Liquibase Helm OpenShift (OCP) CI/CD Automation Infrastructure Deployment Experience with Generative AI and AI-enabled Analytics Platforms. Exposure to Vector Databases, RAG Architectures, Semantic Search, AI Data Platforms. Knowledge of Splunk, Grafana, Observability Platforms. Experience working in Cybersecurity or IAM domains. Kubernetes and Container Platform experience. Exposure to Cloud Cost Optimization and FinOps practices. Benefits: Competitive salary and benefits package Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications Opportunity to work with cutting-edge technologies Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards Annual health check-ups Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents Values-Driven, People-Centric & Inclusive Work Environment: Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds. We support hybrid work and flexible hours to fit diverse lifestyles. Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities. If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment Let’s unleash your full potential at Persistent - persistent.com/careers “Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.”
Requirements
["Java", "Security"]
Roles & responsibilities
[]
About