Source description
About the role
Role Overview We are looking for an experienced SRE Lead with 8+ Yrs of relavant experience to drive reliability, scalability, and observability across our SaaS and AI services. This role will define reliability standards, embed them into CI/CD pipelines, and lead incident response practices to ensure seamless, resilient customer experiences. Key Responsibilities 1. Define and manage Service Level Objectives (SLOs) and Error Budgets for SaaS and AI services. 2. Design and implement the observability architecture, including metrics, logs, and distributed tracing. 3. Lead incident response and conduct blameless postmortems to drive continuous improvement. 4. Champion shiftleft reliability by integrating automated checks into CI/CD pipelines. 5. Collaborate with engineering, QE, and DevOps teams to embed reliability practices across the SDLC. 6. Mentor and guide engineers on best practices in reliability, scalability, and resilience.\ 7. Design and Implement DataDog APM infrastructure metrics, Log management & Dashboards for AKS hosted work loads Required Skills & Experience 1. Proven experience in SRE leadership or senior reliability engineering roles. 2. Solid expertise in CI/CD pipelines, observability tools, and cloud-native architectures. 3. Hands-on experience with monitoring and tracing platforms (e.g., Prometheus, Grafana, ELK, Datadog, OpenTelemetry). 4. Proficiency in infrastructure as code (Terraform, Ansible) and container orchestration (Kubernetes). 5. Solid understanding of SLOs, error budgets, and reliability engineering principles. 6. Excellent leadership, communication, and problem-solving skills. Tools & Platforms 1. CI/CD Pipelines: Jenkins, GitHub Actions, GitLab CI, Azure DevOps 2. Observability & Monitoring: Prometheus, Grafana, ELK/EFK stack, Datadog, New Relic, Splunk, OpenTelemetry 3. Incident Management: PagerDuty, Opsgenie, ServiceNow 4. Cloud Platforms: AWS, Azure, Google Cloud Platform (GCP) 5. Containerization & Orchestration: Docker, Kubernetes, Helm 6. Infrastructure as Code (IaC): Terraform, Ansible, Pulumi, CloudFormation 7. Version Control: Git, GitHub, GitLab, Bitbucket Role Overview We are looking for an experienced SRE Lead with 8+ Yrs of relavant experience to drive reliability, scalability, and observability across our SaaS and AI services. This role will define reliability standards, embed them into CI/CD pipelines, and lead incident response practices to ensure seamless, resilient customer experiences. Key Responsibilities 1. Define and manage Service Level Objectives (SLOs) and Error Budgets for SaaS and AI services. 2. Design and implement the observability architecture, including metrics, logs, and distributed tracing. 3. Lead incident response and conduct blameless postmortems to drive continuous improvement. 4. Champion shiftleft reliability by integrating automated checks into CI/CD pipelines. 5. Collaborate with engineering, QE, and DevOps teams to embed reliability practices across the SDLC. 6. Mentor and guide engineers on best practices in reliability, scalability, and resilience.\ 7. Design and Implement DataDog APM infrastructure metrics, Log management & Dashboards for AKS hosted work loads Required Skills & Experience 1. Proven experience in SRE leadership or senior reliability engineering roles. 2. Solid expertise in CI/CD pipelines, observability tools, and cloud-native architectures. 3. Hands-on experience with monitoring and tracing platforms (e.g., Prometheus, Grafana, ELK, Datadog, OpenTelemetry). 4. Proficiency in infrastructure as code (Terraform, Ansible) and container orchestration (Kubernetes). 5. Solid understanding of SLOs, error budgets, and reliability engineering principles. 6. Excellent leadership, communication, and problem-solving skills. Tools & Platforms 1. CI/CD Pipelines: Jenkins, GitHub Actions, GitLab CI, Azure DevOps 2. Observability & Monitoring: Prometheus, Grafana, ELK/EFK stack, Datadog, New Relic, Splunk, OpenTelemetry 3. Incident Management: PagerDuty, Opsgenie, ServiceNow 4. Cloud Platforms: AWS, Azure, Google Cloud Platform (GCP) 5. Containerization & Orchestration: Docker, Kubernetes, Helm 6. Infrastructure as Code (IaC): Terraform, Ansible, Pulumi, CloudFormation 7. Version Control: Git, GitHub, GitLab, Bitbucket