Source description
About the role
Required Skills:
Strong experience in SRE / Incident Management leadership
Proven ability to manage high-impact, complex incidents
Excellent communication, stakeholder management, and leadership skills
Ability to drive alignment and influence cross-functional teams
Technical Skills:
Ability to proactively identify risks using monitoring tools such as DataDog and Grafana dashboards
Experience in incident response with capability to quickly restore services (restart, patch, or remediate live issues)
Strong focus on minimizing service downtime across environments
Hands-on experience supporting both on-premise (Linux environments) and cloud platforms (primarily Azure, with some exposure to GCP)
Solid understanding of networking concepts and system architecture
More at EXL