Source description
About the role
Position - SR. Site Reliability Engineer Experience - 12+ years Location - Remote for India Employment type- Full time Project- UAE project Skills- Azure, Devops, Python/shell, Job Responsibilities: • Design, implement, and maintain highly reliable, scalable, and observable systems on Microsoft Azure. • Define and operationalize SLIs, SLOs, and SLAs to measure and improve service reliability. • Build and enhance observability platforms using tools like Grafana, Prometheus, ELK stack, and OpenTelemetry. • Drive adoption of SRE principles (error budgets, toil reduction, automation-first mindset). • Implement proactive monitoring, alerting, and incident response frameworks. • Lead incident management, root cause analysis (RCA), and postmortems with a blameless culture. • Automate infrastructure and workflows using Infrastructure as Code (IaC). • Collaborate with engineering teams to improve system resilience, performance, and deployment practices. • Develop and maintain runbooks, playbooks, and operational standards. • Advocate for DevOps and SRE culture adoption across teams Desired Skill: SRE Practices • Hands-on experience implementing: o SLIs, SLOs, SLAs o Error budgets o Toil reduction strategies • Strong understanding of incident management lifecycle Relevant Exp: Programming & Automation • Proficiency in Python (automation, tooling, scripting) • Experience building internal tools for reliability and observability Infrastructure as Code (IaC) • Strong experience with Terraform • Familiarity with infrastructure automation and configuration management • Experience with GitHub Actions (or similar CI/CD tools) • Knowledge of deployment strategies (blue-green, canary, rolling updates) Value Add: Good to Have Experience with AI/ML Observability (monitoring models, drift detection, LLM observability) • Familiarity with: o Service Mesh (Istio, Linkerd) o Chaos Engineering tools (e.g., Chaos Monkey, Litmus) o Distributed tracing tools (Jaeger, Tempo) • Exposure to FinOps practices (cost optimization in cloud) • Experience with multi-cloud or hybrid environments Comment: Technical Skills: Cloud & Infrastructure • Strong hands-on experience with Microsoft Azure • Experience with Kubernetes (AKS) and containerized environments • Knowledge of networking, load balancing, and distributed systems Observability & Monitoring • Experience with: o Grafana (dashboards, alerting) o Prometheus / OpenTelemetry o ELK Stack (Elasticsearch, Logstash, Kibana) • Ability to design end-to-end observability (metrics, logs, traces)
More at Softenger