Padmi

Lead SRE

IndiaPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at JobItUs

Opens the source posting on shine.com

Source description

About the role

View original

As an experienced Lead Site Reliability Engineer (SRE) at our company, you will have the opportunity to lead a dedicated team of 56 engineers in supporting 24x7 production operations. Your expertise in incident management, production support, automation, and cloud infrastructure will be crucial in ensuring the reliability and high availability of our mission-critical systems. Key Responsibilities: - Lead incident response efforts and serve as an Incident Commander when needed - Utilize your strong hands-on experience with AWS services such as EC2, ECS, S3, Lambda, and Load Balancers, as well as Kubernetes - Demonstrate proficient Linux administration and troubleshooting skills to maintain system stability - Implement and enhance CI/CD pipelines using tools like GitHub, GitLab, or Harness - Apply your knowledge of networking concepts like DNS, TCP/IP, HTTP(S), and load balancing to optimize system performance - Troubleshoot Java, Node.js, React applications, and database connectivity issues - Manage hybrid or multi-cloud environments (AWS/Azure/GCP) and automate operational tasks to reduce manual efforts Qualifications Required: - 8+ years of experience in SRE, DevOps, or Production Support - Previous experience in leading incident response and acting as an Incident Commander - Proficiency in AWS services and Kubernetes - Strong Linux administration and troubleshooting skills - Familiarity with CI/CD tools and networking concepts - Experience with Java, Node.js, React applications, and database troubleshooting - Background in hybrid or multi-cloud environments and operational automation In this role, you will be expected to drive reliability, minimize operational toil through automation, mentor team members, and ensure the high availability of our systems in a fast-paced production environment. Your proactive leadership and technical skills will be instrumental in maintaining our systems' performance and stability. As an experienced Lead Site Reliability Engineer (SRE) at our company, you will have the opportunity to lead a dedicated team of 56 engineers in supporting 24x7 production operations. Your expertise in incident management, production support, automation, and cloud infrastructure will be crucial in ensuring the reliability and high availability of our mission-critical systems. Key Responsibilities: - Lead incident response efforts and serve as an Incident Commander when needed - Utilize your strong hands-on experience with AWS services such as EC2, ECS, S3, Lambda, and Load Balancers, as well as Kubernetes - Demonstrate proficient Linux administration and troubleshooting skills to maintain system stability - Implement and enhance CI/CD pipelines using tools like GitHub, GitLab, or Harness - Apply your knowledge of networking concepts like DNS, TCP/IP, HTTP(S), and load balancing to optimize system performance - Troubleshoot Java, Node.js, React applications, and database connectivity issues - Manage hybrid or multi-cloud environments (AWS/Azure/GCP) and automate operational tasks to reduce manual efforts Qualifications Required: - 8+ years of experience in SRE, DevOps, or Production Support - Previous experience in leading incident response and acting as an Incident Commander - Proficiency in AWS services and Kubernetes - Strong Linux administration and troubleshooting skills - Familiarity with CI/CD tools and networking concepts - Experience with Java, Node.js, React applications, and database troubleshooting - Background in hybrid or multi-cloud environments and operational automation In this role, you will be expected to drive reliability, minimize operational toil through automation, mentor team members, and ensure the high availability of our systems in a fast-paced production environment. Your proactive leadership and technical skills will be instrumental in maintaining our systems' performance and stability.

One address, no account. We’ll tell you when matching roles go live.

More at JobItUs

Related open roles

View all roles