← Back to jobs

Data Architect

Persistent Systems · Data_Int_ CRO 1 · Lead

Apply ↗
Location
Pune, Maharashtra, India
Employment type
Full-time
Posted
22-Jul-2026

Skills

["Databricks""PySpark"]

Technology stack

{}

Description

About Persistent We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what’s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem. Our disruptor’s mindset, commitment to client success, and agility to thrive in the dynamic environment have enabled us to sustain our growth momentum. Persistent has been recognized across top industry platforms for innovation, leadership, and inclusion. We reported $1,654.4M FY26 revenue with 17.4% Y-o-Y growth. We have delivered 24 sequential quarters of growth with $436.0M in Q4 FY26 revenue, up 3.2% Q-o-Q and 16.2% Y-o-Y growth. Our 27,500+ global team members, located in 18 countries, have been instrumental in helping the market leaders transform their industries. We have been recognized as the Fastest Growing IT Services Brand Globally in the 2026 Brand Finance IT Services 25 Report. We named a Leader in the Everest Group Private Equity (PE) Services PEAK Matrix® Assessment 2026 and Software Product Engineering PEAK Matrix® Assessment 2026. About Position: We are seeking a highly skilled Data Engineer with strong expertise in Databricks, PySpark, and Azure Cloud to build and maintain scalable enterprise data platforms and analytics solutions. The ideal candidate will be responsible for designing and optimizing ETL/ELT pipelines, implementing Lakehouse architectures, and delivering high-quality, analytics-ready datasets. This role requires hands-on experience with cloud-based data engineering, large-scale data processing, and modern data platform technologies. Role: Data Engineer Location: Pune Experience: 8-12 Years Job Type: Full Time Employment What You'll Do: Design, develop, and maintain scalable ETL/ELT pipelines using Databricks and PySpark. Build and optimize large-scale data ingestion, transformation, and processing workflows. Develop and manage Delta Lake-based Lakehouse architectures and data models. Integrate data from multiple data sources including databases, APIs, cloud storage, and enterprise applications. Implement batch and real-time data processing solutions to support business requirements. Develop and maintain data pipelines using Azure Data Factory (ADF), Azure Data Lake Storage (ADLS), and Azure services. Optimize data processing frameworks for performance, scalability, reliability, and cost efficiency. Collaborate with Data Architects, Business Analysts, Data Scientists, and Analytics teams to deliver business-centric data solutions. Apply data governance, security, compliance, and data quality best practices across the platform. Perform data modeling, warehouse design, partitioning, and optimization activities. Troubleshoot and resolve data pipeline, processing, and performance-related issues. Implement monitoring, alerting, and operational support processes for critical data workloads. Participate in architecture discussions, code reviews, and continuous improvement initiatives. Support CI/CD implementation and automation for data engineering solutions. Expertise You'll Bring: 8 to 12 years of experience in Data Engineering, Data Warehousing, and Cloud Data Platform development. Strong hands-on expertise in Databricks, PySpark, Azure Cloud, SQL, and Python. Proven experience in designing and developing enterprise-scale ETL/ELT solutions. Strong knowledge of Delta Lake and Lakehouse Architecture principles. Experience with Azure Data Factory (ADF), Azure Data Lake Storage (ADLS), and Azure Synapse Analytics. Expertise in data pipeline development, orchestration, and optimization. Solid understanding of data warehousing concepts, dimensional modeling, and data architecture. Experience with batch and streaming data processing frameworks. Strong troubleshooting, performance tuning, and optimization skills. Experience with Apache Airflow for workflow orchestration. Knowledge of Change Data Capture (CDC) methodologies and implementation. Familiarity with data quality frameworks, governance practices, and data security standards. Experience with Git, CI/CD pipelines, and DevOps practices for data platforms. Understanding of monitoring, observability, and alerting frameworks. Strong analytical, problem-solving, and debugging capabilities. Excellent communication, collaboration, and stakeholder management skills. Ability to work independently while driving delivery across cross-functional teams. Values-Driven, People-Centric & Inclusive Work Environment: Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds. We support hybrid work and flexible hours to fit diverse lifestyles. Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities. If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment Let’s unleash your full potential at Persistent - persistent.com/careers “Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.”

Requirements

["Databricks", "PySpark"]

Roles & responsibilities

[]

About

All jobs →