Source description
About the role
Role & responsibilities 8 to 11 years in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering. Experience supporting enterprise-scale production environments. Proven ownership of services with 99.9%+ uptime commitments . Demonstrated experience creating and managing SLOs and Error Budgets. Deep troubleshooting expertise in Kubernetes-based production systems. Experience handling Sev-1 and Sev-2 incidents. Strong understanding of distributed systems and microservices architectures. Hands-on cloud platform administration experience. Defining SLOs & SLIs Have experience in Service mesh Exposure to observability tool Datadog (good to have) Deal Breaker Skill Reliability Analysis Mandatory Skills Site Reliability Engineering| DevOps| Reliability Engineering Desirable Skills Ansible, Reliability Analysis, GitHub Actions, Terraform, Istio, AWS S3
More at Happiest Minds Technologies