← Back to jobs

Lead Gen AI Engineer

Persistent Systems · Data_Int_ Hitech 2 · Lead

Apply ↗
Location
Pune, Maharashtra, India
Employment type
Full-time
Posted
16-Jun-2026

Skills

["AI""SRE""LLM"]

Technology stack

{}

Description

About Persistent We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what?s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem. Our disruptor?s mindset, commitment to client success, and agility to thrive in the dynamic environment have enabled us to sustain our growth momentum. Persistent has been recognized across top industry platforms for innovation, leadership, and inclusion. We reported $1,654.4M FY26 revenue with 17.4% Y-o-Y growth. We have delivered 24 sequential quarters of growth with $436.0M in Q4 FY26 revenue, up 3.2% Q-o-Q and 16.2% Y-o-Y growth. Our 27,500+ global team members, located in 18 countries, have been instrumental in helping the market leaders transform their industries. We have been recognized as the Fastest Growing IT Services Brand Globally in the 2026 Brand Finance IT Services 25 Report. We named a Leader in the Everest Group Private Equity (PE) Services PEAK Matrix? Assessment 2026 and Software Product Engineering PEAK Matrix? Assessment 2026. About Position: We are seeking a Senior AI/GenAI Site Reliability Engineer (SRE) to provide reliability, visibility, and operational governance for AI agent and LLM-based services powering internal productivity. This role ensures consistent telemetry, end-to-end traceability, operational controls, and production readiness across agent systems while bridging engineering delivery and enterprise operational excellence. Role: Site Reliability Engineer Location: Pune Experience: 5 to 8 Years Job Type: Full-Time Employment What You'll Do: Establish and enforce a standardized observability framework across AI agents and LLM-powered services Implement consistent telemetry, correlation IDs, logging, tracing, and monitoring standards Define and track operational metrics related to agent health, performance, reliability, and user experience Build operational controls for AI services, including safe enablement, disablement, rollback, and recovery mechanisms Design actionable alerting and monitoring strategies to proactively identify and resolve issues Create end-to-end traceability frameworks to support debugging, auditing, governance, and compliance requirements Define and implement SRE governance standards for AI and GenAI platforms Establish Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability measurements for AI services Define production onboarding criteria, operational readiness standards, and release governance processes Develop incident response frameworks, escalation procedures, and operational runbooks Partner with engineering teams to ensure services are properly instrumented and production-ready Conduct reliability reviews and operational readiness assessments for AI solutions Drive continuous improvement initiatives focused on service availability, resilience, scalability, and performance Support incident management, root cause analysis, and post-incident reviews Design and mature operational support models, including on-call structures and response processes Contribute to AI-driven operational tooling and automation initiatives Support development of internal AI agents and chatbot solutions to improve SRE productivity and operational efficiency Collaborate with platform, security, architecture, and product teams to ensure enterprise operational standards are met Establish observability dashboards, reporting mechanisms, and operational analytics Promote engineering best practices for reliability, governance, monitoring, and production operations Expertise You'll Bring: Bachelor's degree or higher in Computer Science, Engineering, or a related field 5+ years of experience in Site Reliability Engineering, Production Engineering, Platform Operations, or Infrastructure Operations Strong experience supporting distributed systems and production services at scale Deep understanding of observability principles including monitoring, logging, tracing, and telemetry Experience defining and implementing enterprise-wide observability and operational standards Strong troubleshooting, performance analysis, and incident management capabilities Experience implementing operational controls and service governance models Expertise defining SLIs, SLOs, error budgets, and reliability frameworks Experience designing incident response processes, escalation paths, and operational runbooks Understanding of production readiness reviews and release governance practices Experience supporting cloud-native and microservices-based architectures Knowledge of monitoring and observability platforms such as Datadog, Grafana, Prometheus, New Relic, Splunk, OpenTelemetry, or similar technologies Experience with distributed tracing, correlation IDs, and end-to-end service visibility Exposure to AI, GenAI, LLMs, agent-based systems, or AI-powered applications in production environments Understanding of operational challenges associated with AI/LLM service reliability and governance Experience implementing operational monitoring and compliance controls for enterprise systems Knowledge of automation practices and operational tooling development Understanding of DevOps, Platform Engineering, and SRE methodologies Experience creating dashboards, metrics frameworks, and operational reporting mechanisms Ability to drive operational excellence initiatives across multiple engineering teams Strong analytical, problem-solving, and root cause analysis skills Excellent communication, stakeholder management, and influencing capabilities Ability to define standards and drive adoption across engineering organizations Experience working in Agile and cross-functional environments Strong focus on reliability, governance, scalability, and enterprise operational maturity Experience supporting AI-driven operational tools, chatbots, or automation platforms will be an added advantage Familiarity with enterprise governance, audit, compliance, and security requirements for AI systems will be beneficial Benefits: Competitive salary and benefits package Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications Opportunity to work with cutting-edge technologies Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards Annual health check-ups Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents Values-Driven, People-Centric & Inclusive Work Environment: Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds. We support hybrid work and flexible hours to fit diverse lifestyles. Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities. If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment Let?s unleash your full potential at Persistent - persistent.com/careers ?Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.?

Requirements

["AI", "SRE", "LLM"]

Roles & responsibilities

[]

About

All jobs →