Padmi

Senior Site Reliability Engineer-III

Delhi NCRPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at GREY ORANGE INC

Opens the source posting on shine.com

Source description

About the role

View original

Key Responsibilities: Define and enforce SLOs, SLIs, and error budgets across microservices.Architect an observability stack (metrics, logs, traces) and derive operational insights.Automate toil and manual operations through robust tooling and runbooks.Own the incident response lifecycle: detection, triage, RCA, and postmortems.Collaborate with product teams to build fault-tolerant and scalable systems.Champion performance tuning, capacity planning, and scalability testing.Optimize cloud costs while maintaining reliability of infrastructure.Participate in on-call rotations and manage large-scale production systems. Key Responsibilities: Define and enforce SLOs, SLIs, and error budgets across microservices.Architect an observability stack (metrics, logs, traces) and derive operational insights.Automate toil and manual operations through robust tooling and runbooks.Own the incident response lifecycle: detection, triage, RCA, and postmortems.Collaborate with product teams to build fault-tolerant and scalable systems.Champion performance tuning, capacity planning, and scalability testing.Optimize cloud costs while maintaining reliability of infrastructure.Participate in on-call rotations and manage large-scale production systems.

One address, no account. We’ll tell you when matching roles go live.

More at GREY ORANGE INC

Related open roles

View all roles