Padmi

Lead Site Reliability Engineer - AI Platforms

BangalorePosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at UnitedHealth Group

Opens the source posting on shine.com

Source description

About the role

View original

Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. Primary Responsibilities: Collaborate with research, engineering, and product teams to translate cutting-edge AI advancements into production-ready capabilities. Uphold ethical AI principles by embedding fairness, transparency, and accountability throughout the model development lifecycleComply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do soRequired Qualifications: 8+ years of experience in SRE, DevOps, or Platform Engineering with large-scale systemsHands-on experience with observability, monitoring, logging, tracing, alerting, and production operationsExperience deploying and operating AI inference services, RAG pipelines, vector databases, and AI serving platformsExperience building and supporting CI/CD pipelines, deployment automation, and platform operational workflowsExperience implementing auto-scaling, load balancing, disaster recovery, failover, backup, and business continuity solutionsExperience supporting multi-region, multi-cluster, and distributed cloud environmentsExperience working with event-driven architectures, messaging systems, and real-time processing workloadsExperience optimizing platform performance, resource utilization, AI inference workloads, and operational costsExperience mentoring junior engineers and contributing to engineering best practicesExperience supporting production AI/ML, Generative AI, LLM, or data-intensive platforms Experience with Kubernetes, containerization, and cloud-native deployment practices Experience building and supporting CI/CD pipelines and deployment automation Experience deploying and supporting AI services, APIs, inference endpoints, and RAG-based solutions Experience with Infrastructure as Code (Terraform, CloudFormation, ARM, Pulumi, or equivalent) Experience with monitoring, logging, tracing, observability, and alerting platforms Experience implementing operational controls for backup, recovery, failover, and disaster recovery processes Experience with AWS, Azure, or GCP environments Experience supporting production incidents, troubleshooting, root cause analysis, and operational excellence initiatives Experience optimizing platform reliability, performance, resource utilization, and operational costs Proven experience in SRE, DevOps, Platform Engineering, Cloud Infrastructure, or Production Operations Proven experience supporting and operating production-scale AI/ML, Generative AI, and LLM-based platformsSolid experience implementing MLOps, LLMOps, model deployment, monitoring, and lifecycle management practicesSolid experience with cloud-native technologies, Kubernetes, container orchestration, and Infrastructure as CodeKnowledge of data security, governance, and compliance requirements for enterprise AI platformsKnowledge of cloud security, IAM, RBAC, encryption, secrets management, and security best practices Understanding of distributed systems, scalability, reliability, fault tolerance, and high-availability concepts Good understanding of distributed systems, high availability, scalability, fault tolerance, and reliability engineering principlesGood understanding of security best practices including IAM, RBAC, encryption, secrets management, and Zero Trust principlesFamiliarity with MLOps, LLMOps, model deployment, monitoring, and AI application lifecycle management Familiarity with event-driven architectures, messaging systems, and streaming platforms Solid scripting and automation skills using Python, Bash, PowerShell, or equivalent technologiesSolid scripting and automation skills using Python, Bash, PowerShell, or similar technologies Proven solid troubleshooting, incident management, root cause analysis (RCA), and production support experienceProven ability to independently own platform services and reliability initiatives from implementation through operationsProven solid collaboration and stakeholder management Optum

One address, no account. We’ll tell you when matching roles go live.

More at UnitedHealth Group

Related open roles

View all roles