← Back to jobs

AI Data Platform Engineer - Manufacturing Systems and Infrastructure

Apple · Operations and Supply Chain

Apply ↗
Location
Bengaluru
Employment type
Full-time
Posted
2026-07-21 11:48:04 IST

Skills

["AWS""Airflow""Azure""CI/CD""Docker""GCP""Go""Java""Kafka""Kubernetes""LLMs""Python""RAG""SQL""Scala""Spark"]

Technology stack

{"programming_languages": ["Python""Java""Go""Scala""SQL"]"frameworks": ["Kafka""Spark""Airflow"]"databases": []"cloud": ["AWS""Azure""GCP"]"infrastructure": ["Kubernetes""Docker""CI/CD"]"data_tools": []"ai_ml": ["LLMs""RAG"]"other_tools": []}

Description

As an AI Data Platform Engineer with the MSI team, you will design, build, and operate scalable AI data platforms that enable GenAI, Agentic AI, and Embodied AI solutions across the enterprise. You will develop reusable platform services, data pipelines, and data quality frameworks that transform fragmented enterprise and multimodal data into trusted, AI-ready datasets — combining expertise in AI data platform engineering, data quality, systems engineering, and AI data lifecycle management to accelerate AI innovation.

Requirements

["8+ years of experience designing and building scalable data platforms and distributed systems", "Strong programming skills in Python and SQL, with proficiency in Java or Scala preferred", "Experience with Airflow, Kubeflow, or MLflow to build and orchestrate scalable AI data pipelines", "Experience building scalable batch and streaming data pipelines using Spark (PySpark), Kafka, Airflow, and Ray, with proficiency in Pandas and modern data lake/lakehouse architectures (e.g., Iceberg, Delta Lake)", "Hands-on experience with AI data engineering, including ground truth dataset creation, data curation, annotation pipelines, dataset versioning, and metadata management", "Bachelors / Masters in Computer Science or related fields", "Preferred: Experience implementing data validation, quality frameworks, observability, and AI dataset evaluation", "Preferred: Knowledge of RAG architectures, embedding generation, vector databases, and AI data preparation for LLMs and agentic AI", "Preferred: Experience with cloud platforms (AWS, Azure, or GCP), Kubernetes, Docker, CI/CD, and Infrastructure as Code", "Preferred: Strong understanding of distributed systems, APIs, microservices, and enterprise integration patterns", "Preferred: Excellent communication, collaboration, and technical leadership skills", "Preferred: Experience building platforms supporting GenAI, Agentic AI, or Embodied AI applications", "Preferred: Experience with multimodal datasets, knowledge graphs, AI evaluation frameworks, or vector search technologies", "Preferred: Familiarity with enterprise data governance, lineage, metadata management, and AI compliance", "Preferred: Experience working with manufacturing, operational, IoT, or industrial data platforms", "Preferred: Demonstrated ability to lead technical initiatives and mentor engineers"]

Roles & responsibilities

[]

About

All jobs →