Source description
About the role
Lead Site Reliability Engineer (SRE) We are seeking an experienced Lead Site Reliability Engineer (SRE) to lead a team of 56 engineers supporting 24x7 production operations. The ideal candidate will have strong expertise in incident management, production support, automation, and cloud infrastructure. Key Requirements: 8+ years of experience in SRE, DevOps, or Production SupportProven experience leading incident response and acting as an Incident CommanderStrong hands-on experience with AWS (EC2, ECS, S3, Lambda, Load Balancers) and KubernetesSolid Linux administration and troubleshooting skillsExperience with CI/CD tools such as GitHub, GitLab, or HarnessKnowledge of networking concepts including DNS, TCP/IP, HTTP(S), and load balancingFamiliarity with Java, Node.js, React applications, and database connectivity troubleshootingExperience in hybrid or multi-cloud environments (AWS/Azure/GCP) and operational automationThe role requires a proactive leader who can drive reliability, reduce operational toil through automation, mentor team members, and ensure high availability of mission-critical systems in a fast-paced production environment. Skills Required Github, React, Kubernetes, Gitlab, Gcp, Aws Ec2, Load Balancers, Linux Administration, Dns, Harness, Http, Azure, Java, Node.js, Aws Lead Site Reliability Engineer (SRE) We are seeking an experienced Lead Site Reliability Engineer (SRE) to lead a team of 56 engineers supporting 24x7 production operations. The ideal candidate will have strong expertise in incident management, production support, automation, and cloud infrastructure. Key Requirements: 8+ years of experience in SRE, DevOps, or Production SupportProven experience leading incident response and acting as an Incident CommanderStrong hands-on experience with AWS (EC2, ECS, S3, Lambda, Load Balancers) and KubernetesSolid Linux administration and troubleshooting skillsExperience with CI/CD tools such as GitHub, GitLab, or HarnessKnowledge of networking concepts including DNS, TCP/IP, HTTP(S), and load balancingFamiliarity with Java, Node.js, React applications, and database connectivity troubleshootingExperience in hybrid or multi-cloud environments (AWS/Azure/GCP) and operational automationThe role requires a proactive leader who can drive reliability, reduce operational toil through automation, mentor team members, and ensure high availability of mission-critical systems in a fast-paced production environment. Skills Required Github, React, Kubernetes, Gitlab, Gcp, Aws Ec2, Load Balancers, Linux Administration, Dns, Harness, Http, Azure, Java, Node.js, Aws
More at MM Management Consultant