Padmi

Infrastructure Associate Advisor

HyderabadPosted 3 months ago
Software engineeringSeniorFull Time
Apply at Evernorth Health Services

Opens the source posting on foundit.in

Source description

About the role

View original

Infrastructure Engineering Associate Advisor Position Overvie wThe Pharmacy Benefit Services+ Technology organization is seeking a Site Reliability Engineer (SRE) – Automation, SelfHealing & AI/AIOps to join our team. This Band 4 Contributor role is a senior, handson position responsible for driving enterprise reliability outcomes, reducing operational toil, and enabling scalable SRE adoption across both legacy platforms and modern cloudnative systems .In this role, you will lead the design and implementation of intelligent, automated, and AIassisted reliability solutions that ensure systems are resilient, observable, selfhealing, and continuously improving. You will operate at the intersection of software engineering, operations, automation, and AI, influencing how teams design, deploy, and operate production systems .A core focus of this role is building automationfirst and agentic SRE capabilities, including :Selfhealing workflows that automatically detect, diagnose, and remediate failure sAIdriven operational intelligence (AIOps) for anomaly detection, alert correlation, incident triage, and guided remediatio nStandardized SRE enablement platforms (SLO automation, reliability scorecards, FMEA workflows) that can be adopted at scale with minimal frictio n You will collaborate closely with application teams, platform engineering, DevOps, infrastructure, QE, and IT leadership to embed reliability into the SDLC and runtime operation s.Your contributions will directly suppor t:Improved system availability and resilience through proactive reliability engineering and automati onReduced incidents and faster MTTR via selfhealing and AIassisted operatio nsHigher developer productivity by eliminating manual operational to ilFaster, safer releases by integrating SRE controls into CI/CD pipelin esMeasurable reliability improvements, including reductions in MTTD/MTTR, decreased incident frequency, improved SLO compliance, and healthier errorbudget consumption through automation and AIenabled operatio ns Section 3: Responsibilit iesClearly outline the primary duties and tasks associated with the role. Use action verbs (i.e., lead, drive, analyze, assess, research, etc.) to convey expectatio ns. Responsibilit ies Core SRE & Reliability Enginee ringDefine, implement, and operationalize SRE best practices including SLIs, SLOs, error budgets, reliability reviews, and operational readiness standards across multiple teams and platfo rms.Act as a senior reliability engineer for missioncritical systems, influencing architectural decisions to improve availability, scalability, and fault tolera nce.Lead blameless incident response, rootcause analysis, and postincident reviews, ensuring systemic fixes and automation are prioriti zed.SelfHealing Automa tionDesign and implement selfhealing systems t hat:Automatically detect failures using telemetry and sig nalsDiagnose probable root causes using rules and AI/ML mo delsExecute automated remediation actions (restart, scale, reroute, rollback, configuration correct ion)Build eventdriven automation workflows integrated with monitoring, CI/CD, and infrastructure platforms to reduce human intervent ion.Develop and maintain automated runbooks and remediation pipelines that evolve based on historical incidents and outco mes.Extend selfhealing automation to change and release workflows, including automated rollback, safeguarded change execution, and AIassisted changerisk evaluation to reduce changerelated incide nts.Continuously identify and eliminate operational toil by replacing repetitive manual work with automation, selfservice tooling, and intelligent remediat ion.AI / Agentic SRE (AI Ops)Apply AI and ML techniques to improve operational intelligence, includ ing:Anomaly detection across metrics, logs, and tr acesAlert deduplication, correlation, and noise reduc tionIntelligent incident summarization and impact anal ysisPredictive failure detection and capacityrisk forecas tingImplement agentic AI patterns where autonomous or semiautonomous age nts:Continuously monitor system he althPropose or execute remediation act ionsLearn from past incidents and operator feed backEstablish continuous learning feedback loops to tune AI models and agent behavior based on false positives, incident outcomes, and operator rev iew.Partner with platform and security teams to ensure responsible, secure, and compliant use of AI in production operati ons.Observability & Telem etryEstablish observabilitybydesign standards for applications and platforms (metrics, logs, traces, even ts).Improve signal quality and alerting strategies to focus on user and business impact, not infrastructure no ise.Build and maintain reliability dashboards and scorecards that provide realtime and historical insights into service hea lth.Resilience, Performance & Valida tionDrive resilience and faulttolerance validation using chaos engineering and controlled failure inject ion.Partner with performance and platform teams to ensure systems meet performance, scalability, and recovery objecti ves.Promote safe testing practices for legacyintegrated systems (e.g., service virtualization where direct backend calls pose ri sk). CI/CD & Platform Enabl ementEmbed SRE controls into CI/CD pipelines, inclu ding:SLO validation gatesAutomated canary ana lysisRelease health checks and rollback tri ggersBuild reusable SRE platforms, templates, and onboarding kits that enable teams to adopt reliability practices with minimal manual ef fort.Mentor engineers and act as a technical leader for SRE adoption across the organiza tion.Section 4: Qualifica tionsSpecify the skills, experience, and education required for the role. Differentiate between the must-haves and nice-to-ha ves.Required s kills: List the specific skills required for the job, including technical, leadership skills, and any industry-specific sk ills.Required Exper ience: Clearly state any mandatory requirements, such as formal education, certifications, licenses, or specific years of experi ence.Desired Exper ience: List any nice-to-have experience, including industry experience, exposure to specific technologies, certifications, etc. Qualific ations Required Skills:Site Reliability Engi neering: Deep handson experience with SLOs, error budgets, incident management, and production oper ations.Automation & Software Engi neering: Strong development ski lls in Python, Go, Java, or similar, with the ability to build productiongrade automation and se rvices.SelfHealing Systems: Proven experience designing and implementing automated remediation and closedloop recovery wor kflows.AI / AIOps: Experience applying AI/ML to operations, such as anomaly detection, alert correlation, predictive analysis, or intelligent remed iation.Observ ability: Expertise with platforms s uch as Dynatrace, Prometheus, Grafana, Splunk, AppD ynamics, or equi valent.Cloud & Distributed Systems: Strong understand ing of AWS / Azur e / GCP, microservice s, and Kubernetes / Op e nShift.CI/CD & DevOps: Experience integrating reliability checks and automation into delivery pip elines.Infrastructure as Code: Terraform, CloudFormation, or s imilar.Legacy + Modern Engi neering: Ability to support and modernize reliability practices across monoliths, batch jobs, messaging, and mainframeintegrated s ystems.Leadership & In fluence: Ability to lead through influence, mentor others, and drive adoption across multiple teams. Required Experience & Ed ucation:Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent expe rience). 7+ years of experience in SRE, DevOps, platform engineering, or production software engineerin g roles.Demonstrated success del ivering enterprisescale automation, selfhealing, and reliability impr o vements. Desired Exp erience: Experience building or contrib uting to enterprise SRE enablement platforms (SLO automation, reliability scorecards, FMEA wo rkflows).Handson experie nce with chaos en gineering and resilience testing in productionlike envi ronments.Familiar ity with ServiceNow / CMDB / service modeling to support operational readiness and dependency vi sibility.Experience applying Gene rative AI for operational use cases such as runbook generation, incident summarization, and knowledge r etrieval.Demonstrated del ivery of quantifiable reliability imp rovements (e.g., MTTR reduction, incident volume reduction, improved SLO ad herence).Experience mentoring engineers and sh aping an automationfirst, reliabilitydrive n culture. These two sections will be standardized in the JD template and made not editable. Location & Ho urs of WorkFull-time position, working 40 hours per week. Expected overlap with US hours as appropriatePrimarily based in the Innovation Hub in Hyderabad, India in a hybrid working model (3 days WFO and 2 days WAH)

One address, no account. We’ll tell you when matching roles go live.

More at Evernorth Health Services

Related open roles

View all roles
Infrastructure Associate Advisor at Evernorth Health Services · Padmi