Source description
About the role
We are looking for experienced Azure Site Reliability Engineers (SREs) to support and enhance the reliability, availability, and performance of mission-critical banking systems. Location: Hyderabad Experience: 712 Years Required Technical Expertise Cloud & Infrastructure Microsoft AzureKubernetesOpenShift Observability & Monitoring DatadogDynatrace and/or AppDynamicsSplunk Automation & CI/CD JenkinsAnsiblePython Messaging & Application Technologies KafkaRabbitMQExposure to Java and Node.js Production Support & Incident Management Strong experience handling major production incidentsExperience supporting high-availability, mission-critical environmentsStrong expertise in Root Cause Analysis (RCA) and implementing long-term reliability improvements Key Responsibilities Engineer and enhance observability across systems and platforms Define, implement, and track SLIs and SLOs Build automation for recovery and self-healing Implement cloud-native resiliency and failure-isolation patterns Lead major incident response with an engineering-driven approach Drive system-level root cause fixes Reduce long-term incident volume through reliability engineering initiatives Analyze and optimize CI/CD pipelines to improve reliability outcomes Good to Have: Experience using AIOps for predictive reliability insights. Important: We are looking for genuine SRE profiles with strong Azure, observability, automation, and major incident management experience. We are looking for experienced Azure Site Reliability Engineers (SREs) to support and enhance the reliability, availability, and performance of mission-critical banking systems. Location: Hyderabad Experience: 712 Years Required Technical Expertise Cloud & Infrastructure Microsoft AzureKubernetesOpenShift Observability & Monitoring DatadogDynatrace and/or AppDynamicsSplunk Automation & CI/CD JenkinsAnsiblePython Messaging & Application Technologies KafkaRabbitMQExposure to Java and Node.js Production Support & Incident Management Strong experience handling major production incidentsExperience supporting high-availability, mission-critical environmentsStrong expertise in Root Cause Analysis (RCA) and implementing long-term reliability improvements Key Responsibilities Engineer and enhance observability across systems and platforms Define, implement, and track SLIs and SLOs Build automation for recovery and self-healing Implement cloud-native resiliency and failure-isolation patterns Lead major incident response with an engineering-driven approach Drive system-level root cause fixes Reduce long-term incident volume through reliability engineering initiatives Analyze and optimize CI/CD pipelines to improve reliability outcomes Good to Have: Experience using AIOps for predictive reliability insights. Important: We are looking for genuine SRE profiles with strong Azure, observability, automation, and major incident management experience.
More at Principle Pride