Source description
About the role
As a proactive and detail-oriented Site Reliability Engineer (SRE) with 3+ years of experience, your role will focus on ensuring high availability, reliability, and performance of production systems through automation, incident management, and cross-team coordination. Your key responsibilities will include: - Maintaining reliable, scalable, and secure production environments. - Implementing and managing monitoring, alerting, and logging solutions. - Contributing to defining and tracking SLIs/SLOs and supporting error budget practices. - Automating operational tasks to improve efficiency and reduce manual effort. - Performing troubleshooting and Root Cause Analysis (RCA) for production incidents. - Optimizing system performance, availability, and capacity. - Maintaining SOPs and incident documentation in Confluence. - Adhering to change management, deployment governance, and disaster recovery standards. - Supporting incident response for critical production services. In terms of Collaboration & Tools, you will be expected to: - Coordinate with external vendors and internal cross-functional teams. - Work closely with Engineering, Product Owners, and Operations teams. - Manage incidents and changes using ServiceNow & JIRA. - Collaborate through Slack and structured communication channels. Your technical skills in Systems & Clouds should include: - Strong knowledge of Windows and Linux/Unix systems. - Solid understanding of networking fundamentals (DNS, TCP/IP, Load Balancing, Firewalls). - Experience with at least one cloud platform (AWS, Azure, or GCP). Additional technical skills required are in Automation & CI/CD, Containers, and ITSM & Documentation. Preferred additional experience includes background in DevOps, Cloud Engineering, or Platform Engineering, understanding of security best practices and compliance standards, familiarity with AI-assisted engineering tools, and exposure to large-scale or production-grade systems. Moreover, your soft skills should include: - Strong analytical and troubleshooting mindset. - Excellent written and verbal communication skills. - Ownership driven and composed during high-level severity incidents. The company is committed to creating an inclusive environment for all employees, including persons with disabilities, and reasonable accommodations will be provided upon request. As a proactive and detail-oriented Site Reliability Engineer (SRE) with 3+ years of experience, your role will focus on ensuring high availability, reliability, and performance of production systems through automation, incident management, and cross-team coordination. Your key responsibilities will include: - Maintaining reliable, scalable, and secure production environments. - Implementing and managing monitoring, alerting, and logging solutions. - Contributing to defining and tracking SLIs/SLOs and supporting error budget practices. - Automating operational tasks to improve efficiency and reduce manual effort. - Performing troubleshooting and Root Cause Analysis (RCA) for production incidents. - Optimizing system performance, availability, and capacity. - Maintaining SOPs and incident documentation in Confluence. - Adhering to change management, deployment governance, and disaster recovery standards. - Supporting incident response for critical production services. In terms of Collaboration & Tools, you will be expected to: - Coordinate with external vendors and internal cross-functional teams. - Work closely with Engineering, Product Owners, and Operations teams. - Manage incidents and changes using ServiceNow & JIRA. - Collaborate through Slack and structured communication channels. Your technical skills in Systems & Clouds should include: - Strong knowledge of Windows and Linux/Unix systems. - Solid understanding of networking fundamentals (DNS, TCP/IP, Load Balancing, Firewalls). - Experience with at least one cloud platform (AWS, Azure, or GCP). Additional technical skills required are in Automation & CI/CD, Containers, and ITSM & Documentation. Preferred additional experience includes background in DevOps, Cloud Engineering, or Platform Engineering, understanding of security best practices and compliance standards, familiarity with AI-assisted engineering tools, and exposure to large-scale or production-grade systems. Moreover, your soft skills should include: - Strong analytical and troubleshooting mindset. - Excellent written and verbal communication skills. - Ownership driven and composed during high-level severity incidents. The company is committed to creating an inclusive environment for all employees, including persons with disabilities, and reasonable accommodations will be provided upon request.
More at Right Advisors
Related open roles
Full Stack Developers React Js, Java, Node Js, Typescript Faridabad
India
Full Stack Developers- (React.js, Java, Node.js, Typescript) (Faridabad)
India
Director Engineering-Backend
Bangalore
Sr. Principal Software Eng.
Hyderabad
Full Stack Developers- (React.js, Java, Node.js, Typescript)
India
Asp.net Developer
India