Padmi

Senior Site Reliability Engineer

IndiaPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time
Apply at Softenger

Opens the source posting on foundit.in

Source description

About the role

View original

Position - SR. Site Reliability Engineer Experience - 12+ years Location - Remote for India Employment type- Full time Project- UAE project Skills- Azure, Devops, Python/shell, Job Responsibilities: • Design, implement, and maintain highly reliable, scalable, and observable systems on Microsoft Azure. • Define and operationalize SLIs, SLOs, and SLAs to measure and improve service reliability. • Build and enhance observability platforms using tools like Grafana, Prometheus, ELK stack, and OpenTelemetry. • Drive adoption of SRE principles (error budgets, toil reduction, automation-first mindset). • Implement proactive monitoring, alerting, and incident response frameworks. • Lead incident management, root cause analysis (RCA), and postmortems with a blameless culture. • Automate infrastructure and workflows using Infrastructure as Code (IaC). • Collaborate with engineering teams to improve system resilience, performance, and deployment practices. • Develop and maintain runbooks, playbooks, and operational standards. • Advocate for DevOps and SRE culture adoption across teams Desired Skill: SRE Practices • Hands-on experience implementing: o SLIs, SLOs, SLAs o Error budgets o Toil reduction strategies • Strong understanding of incident management lifecycle Relevant Exp: Programming & Automation • Proficiency in Python (automation, tooling, scripting) • Experience building internal tools for reliability and observability Infrastructure as Code (IaC) • Strong experience with Terraform • Familiarity with infrastructure automation and configuration management • Experience with GitHub Actions (or similar CI/CD tools) • Knowledge of deployment strategies (blue-green, canary, rolling updates) Value Add: Good to Have Experience with AI/ML Observability (monitoring models, drift detection, LLM observability) • Familiarity with: o Service Mesh (Istio, Linkerd) o Chaos Engineering tools (e.g., Chaos Monkey, Litmus) o Distributed tracing tools (Jaeger, Tempo) • Exposure to FinOps practices (cost optimization in cloud) • Experience with multi-cloud or hybrid environments Comment: Technical Skills: Cloud & Infrastructure • Strong hands-on experience with Microsoft Azure • Experience with Kubernetes (AKS) and containerized environments • Knowledge of networking, load balancing, and distributed systems Observability & Monitoring • Experience with: o Grafana (dashboards, alerting) o Prometheus / OpenTelemetry o ELK Stack (Elasticsearch, Logstash, Kibana) • Ability to design end-to-end observability (metrics, logs, traces)

One address, no account. We’ll tell you when matching roles go live.

More at Softenger

Related open roles

View all roles
Senior Site Reliability Engineer at Softenger · Padmi