Source description
About the role
Location :- Hybrid /Banglore Time :- 7:30 AM-5 PM """ Summary A Site Reliability Engineer (SRE) applies software engineering practices to operations to ensure services are reliable, scalable, and efficient. The SRE partner with product and platform teams to define SLIs/SLOs, automate operations, runbook and incident engineering, and drive long term reliability improvements. Core Responsibilities Service Reliability: Define SLIs/SLOs and monitor error budgets; drive actions when SLOs are at risk. Incident Management: Lead on call rotations, perform incident response, run postmortems, and implement corrective actions. Automation & Tooling: Build automation for deployment, remediation, scaling, and runbook tasks to reduce manual toil. Observability: Design and maintain metrics, logs, and distributed tracing to support rapid diagnosis and capacity planning. Performance & Capacity: Run load tests, capacity planning, and tuning to meet performance targets. Resilience Engineering: Implement canaries, chaos testing, circuit breakers, and progressive rollouts. Platform Improvement: Collaborate with dev teams to productionize features, harden services, and reduce operational risk. Knowledge Sharing: Produce runbooks, run regular reliability reviews, and coach teams on best practices. Required Skills & Experience Core: Strong programming/scripting (Python, Go, or equivalent), systems fundamentals (Linux), networking, and debugging skills. Observability: Experience with metrics platforms, logging, and tracing (Prometheus, Grafana, ELK/Opentelemetry or equivalent). Cloud & Infra: Familiarity with cloud platforms (AWS/Azure/GCP), containers, orchestration (Kubernetes), and IaC (Terraform/ARM). Operational Practice: Proven incident leadership, postmortem facilitation, and on call experience. Automation: CI/CD pipelines, deployment automation, and infrastructure automation experience. Soft skills: Clear communicator, calm under pressure, collaborative, and able to drive cross team change. Preferred Qualifications Bachelor s in computer science, Engineering, or equivalent experience. Experience defining SLIs/SLOs and using error budget processes. Prior SRE, production operations, or systems engineering role. Certifications or courses in cloud, Kubernetes, or reliability engineering (optional). Success Metrics (KPIs) Achieve and sustain defined SLO targets for owned services. Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR) improvements quarter over quarter. Reduction in manual toil hours via automation. Number and quality of postmortem action items closed within SLA. Error budget burn rate and governance adherence.""" Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at Sparix Global