Source description
About the role
SRE Lead, Platform Reliability and Operations - NewLeaf Azure application platform This is the role specification for a senior reliability leader who can combine SRE discipline, platform operations, and automation-first thinking across the NewLeaf Azure application estate. Role Summary We are hiring a hands-on SRE Lead to own reliability across the full NewLeaf application platform on Azure. This role covers customer-facing websites, backend services, APIs, integrations, event-driven workflows, observability, incident management, self-healing automation, on-call operations, and operational reporting. The goal is to build a resilient operating model for the entire product ecosystem, not just the underlying cloud infrastructure. What "Platform" Means Here Public and internal websites Backend services and APIs Azure-hosted application workloads Integrations and event processing Monitoring, alerting, dashboards, and reporting Incident response, postmortems, and reliability improvements Runbooks, operational standards, and self-healing automation On-call rotation design and engineering readiness Primary Responsibilities Define the reliability strategy for the NewLeaf platform and keep it aligned with business priorities. Own incident detection, triage, escalation, communication, and post-incident follow-through. Build operational dashboards that show service health, user impact, and reliability trends. Design self-healing and auto-remediation workflows for recurring failure modes. Create and maintain runbooks, playbooks, and a living reliability rulebook. Design and support a sustainable on-call rotation model for engineering teams. Drive root-cause analysis and convert recurring issues into permanent fixes or automation. Partner with application, integration, and platform teams to reduce operational risk across boundaries. Required Skill Set Solid SRE or platform engineering background with direct ownership of production systems. D .
More at Qualitest