Source description
About the role
As an experienced Site Reliability Engineering (SRE) Manager, your role will involve leading and managing a team of SRE/DevOps engineers. You will be responsible for defining and implementing SRE best practices such as SLIs, SLOs, and error budgets. Your expertise will ensure system reliability, scalability, and performance while driving automation initiatives and collaborating with cross-functional teams. Additionally, you will own CI/CD pipelines, lead incident response and RCA processes, establish monitoring frameworks, and manage cloud infrastructure across AWS, Azure, and GCP. Your role will also involve implementing disaster recovery plans. Key Responsibilities: - Lead, mentor, and manage a team of SRE/DevOps engineers - Define and implement SRE best practices (SLIs, SLOs, error budgets) - Ensure system reliability, scalability, and performance - Drive automation initiatives - Collaborate with cross-functional teams - Own CI/CD pipelines and release management - Lead incident response and RCA processes - Establish monitoring and observability frameworks - Manage cloud infrastructure (AWS/Azure/GCP) - Implement disaster recovery plans Required Skills & Qualifications: - 7+ years of experience in SRE/DevOps roles - 3+ years of team management experience - Experience with cloud platforms (AWS/Azure/GCP) - Knowledge of CI/CD tools (Jenkins, GitLab CI) - Experience with Docker and Kubernetes - Scripting skills (Python, Bash) - Knowledge of Terraform/CloudFormation - Monitoring tools (Prometheus, Grafana, ELK) Preferred Qualifications: - Experience with microservices - Cloud certifications are a plus - Strong problem-solving skills Key Competencies: - Leadership - Communication - Ownership - Stakeholder management Good to Have: - Experience in e-commerce platforms - Knowledge of chaos engineering As an experienced Site Reliability Engineering (SRE) Manager, your role will involve leading and managing a team of SRE/DevOps engineers. You will be responsible for defining and implementing SRE best practices such as SLIs, SLOs, and error budgets. Your expertise will ensure system reliability, scalability, and performance while driving automation initiatives and collaborating with cross-functional teams. Additionally, you will own CI/CD pipelines, lead incident response and RCA processes, establish monitoring frameworks, and manage cloud infrastructure across AWS, Azure, and GCP. Your role will also involve implementing disaster recovery plans. Key Responsibilities: - Lead, mentor, and manage a team of SRE/DevOps engineers - Define and implement SRE best practices (SLIs, SLOs, error budgets) - Ensure system reliability, scalability, and performance - Drive automation initiatives - Collaborate with cross-functional teams - Own CI/CD pipelines and release management - Lead incident response and RCA processes - Establish monitoring and observability frameworks - Manage cloud infrastructure (AWS/Azure/GCP) - Implement disaster recovery plans Required Skills & Qualifications: - 7+ years of experience in SRE/DevOps roles - 3+ years of team management experience - Experience with cloud platforms (AWS/Azure/GCP) - Knowledge of CI/CD tools (Jenkins, GitLab CI) - Experience with Docker and Kubernetes - Scripting skills (Python, Bash) - Knowledge of Terraform/CloudFormation - Monitoring tools (Prometheus, Grafana, ELK) Preferred Qualifications: - Experience with microservices - Cloud certifications are a plus - Strong problem-solving skills Key Competencies: - Leadership - Communication - Ownership - Stakeholder management Good to Have: - Experience in e-commerce platforms - Knowledge of chaos engineering