Padmi

Urgent Hiring | Azure Site Reliability Engineer (SRE)

HyderabadPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Principle Pride

Opens the source posting on shine.com

Source description

About the role

View original

We are looking for experienced Azure Site Reliability Engineers (SREs) to support and enhance the reliability, availability, and performance of mission-critical banking systems. Location: Hyderabad Experience: 712 Years Required Technical Expertise Cloud & Infrastructure Microsoft AzureKubernetesOpenShift Observability & Monitoring DatadogDynatrace and/or AppDynamicsSplunk Automation & CI/CD JenkinsAnsiblePython Messaging & Application Technologies KafkaRabbitMQExposure to Java and Node.js Production Support & Incident Management Strong experience handling major production incidentsExperience supporting high-availability, mission-critical environmentsStrong expertise in Root Cause Analysis (RCA) and implementing long-term reliability improvements Key Responsibilities Engineer and enhance observability across systems and platforms Define, implement, and track SLIs and SLOs Build automation for recovery and self-healing Implement cloud-native resiliency and failure-isolation patterns Lead major incident response with an engineering-driven approach Drive system-level root cause fixes Reduce long-term incident volume through reliability engineering initiatives Analyze and optimize CI/CD pipelines to improve reliability outcomes Good to Have: Experience using AIOps for predictive reliability insights. Important: We are looking for genuine SRE profiles with strong Azure, observability, automation, and major incident management experience. We are looking for experienced Azure Site Reliability Engineers (SREs) to support and enhance the reliability, availability, and performance of mission-critical banking systems. Location: Hyderabad Experience: 712 Years Required Technical Expertise Cloud & Infrastructure Microsoft AzureKubernetesOpenShift Observability & Monitoring DatadogDynatrace and/or AppDynamicsSplunk Automation & CI/CD JenkinsAnsiblePython Messaging & Application Technologies KafkaRabbitMQExposure to Java and Node.js Production Support & Incident Management Strong experience handling major production incidentsExperience supporting high-availability, mission-critical environmentsStrong expertise in Root Cause Analysis (RCA) and implementing long-term reliability improvements Key Responsibilities Engineer and enhance observability across systems and platforms Define, implement, and track SLIs and SLOs Build automation for recovery and self-healing Implement cloud-native resiliency and failure-isolation patterns Lead major incident response with an engineering-driven approach Drive system-level root cause fixes Reduce long-term incident volume through reliability engineering initiatives Analyze and optimize CI/CD pipelines to improve reliability outcomes Good to Have: Experience using AIOps for predictive reliability insights. Important: We are looking for genuine SRE profiles with strong Azure, observability, automation, and major incident management experience.

One address, no account. We’ll tell you when matching roles go live.

More at Principle Pride

Related open roles

View all roles