Padmi

Manager, Site Reliability Engineering

Delhi NCRPosted 3 months ago
Technology ManagementSeniorFull Time; Regular
Apply at Uplers

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: You are required to lead as a Site Reliability Engineering (SRE) Manager for Monotype, ensuring the reliability, stability, and operational excellence of enterprise platforms. Your responsibilities will include owning incident management operations, driving automation, monitoring, and scalability of AI-driven systems. Key Responsibilities: - Own end-to-end reliability of production systems, ensuring uptime within defined SLAs - Lead and govern a 24x7x365 incident management team for quick response and resolution - Drive a blameless RCA culture, analyze root causes, and track closure of action items - Improve observability using tools like Datadog, CloudWatch, ELK, Prometheus - Drive automation to reduce manual effort and operational toil - Collaborate with Product, Engineering & Platform teams to improve release quality and stability - Support reliability and monitoring of AI/ML workloads in production and experimentation environments - Lead and mentor a team of ~14 engineers across operations and SRE excellence - Partner with teams to optimize cloud usage and reduce unnecessary spend - Ensure security best practices are followed across infrastructure and applications Qualifications Required: - Bachelors degree in computer science, Engineering, or related field - 10+ years of experience in SRE managing production systems and operations teams - Strong hands-on experience with AWS and Kubernetes (EKS preferred) - Experience with incident management, RCA, monitoring tools, automation, and release processes - Understanding of microservices-based architectures and cloud cost optimization - Exposure to supporting AI/ML workloads and leadership skills - Certification in relevant technologies (e.g., AWS, Kubernetes) is a plus - Strong analytical, problem-solving, and communication skills Additional Details: Monotype, a global leader in fonts, values reliability, innovation, and collaboration. Monotype Solutions India is a certified Great Place to Work, focusing on various areas like Product Development, User Research, AI, and Machine learning. Monotype aims to bring brands to life through type and technology, providing font solutions for creative professionals worldwide. As an employee, you can expect a creative, innovative work environment with opportunities for career advancement and personal growth. If you are ready for a new challenge and wish to contribute to a global brand, apply now and join the Monotype team! Role Overview: You are required to lead as a Site Reliability Engineering (SRE) Manager for Monotype, ensuring the reliability, stability, and operational excellence of enterprise platforms. Your responsibilities will include owning incident management operations, driving automation, monitoring, and scalability of AI-driven systems. Key Responsibilities: - Own end-to-end reliability of production systems, ensuring uptime within defined SLAs - Lead and govern a 24x7x365 incident management team for quick response and resolution - Drive a blameless RCA culture, analyze root causes, and track closure of action items - Improve observability using tools like Datadog, CloudWatch, ELK, Prometheus - Drive automation to reduce manual effort and operational toil - Collaborate with Product, Engineering & Platform teams to improve release quality and stability - Support reliability and monitoring of AI/ML workloads in production and experimentation environments - Lead and mentor a team of ~14 engineers across operations and SRE excellence - Partner with teams to optimize cloud usage and reduce unnecessary spend - Ensure security best practices are followed across infrastructure and applications Qualifications Required: - Bachelors degree in computer science, Engineering, or related field - 10+ years of experience in SRE managing production systems and operations teams - Strong hands-on experience with AWS and Kubernetes (EKS preferred) - Experience with incident management, RCA, monitoring tools, automation, and release processes - Understanding of microservices-based architectures and cloud cost optimization - Exposure to supporting AI/ML workloads and leadership skills - Certification in relevant technologies (e.g., AWS, Kubernetes) is a plus - Strong analytical, problem-solving, and communication skills Additional Details: Monotype, a global leader in fonts, values reliability, innovation, and collaboration. Monotype Solutions India is a certified Great Place to Work, focusing on various areas like Product Development, User Research, AI, and Machine learning. Monotype aims to bring brands to life through type and technology, providing font solutions for creative professionals worldwide. As an employee, you can expect a creative, innovative work environment with opportunities for career advancement and personal growth. If you are ready for a new challenge and wish to contribute to a global brand, apply now and join the Monotype team!

One address, no account. We’ll tell you when matching roles go live.

More at Uplers

Related open roles

View all roles