Source description
About the role
Experience 3+ years in an SRE roleProven track record in an incident responseExperience supporting production services at scale on a major cloud provider (Azure, AWS, or GCP)Experience working with API gateway or proxy infrastructure is a plus Must-Have Skills Prometheus & Grafana writing PromQL queries from scratch, building dashboards, and interpreting resultsIncident communication structured, calm, and clear written and verbal updates under pressureObservability fundamentals metrics, logs, and traces; understanding the difference and when to use eachCloud fluency working knowledge of at least one major provider's managed services and common failure modesDocumentation ability to write clear runbooks, incident reports, and post-mortems Experience 3+ years in an SRE roleProven track record in an incident responseExperience supporting production services at scale on a major cloud provider (Azure, AWS, or GCP)Experience working with API gateway or proxy infrastructure is a plus Must-Have Skills Prometheus & Grafana writing PromQL queries from scratch, building dashboards, and interpreting resultsIncident communication structured, calm, and clear written and verbal updates under pressureObservability fundamentals metrics, logs, and traces; understanding the difference and when to use eachCloud fluency working knowledge of at least one major provider's managed services and common failure modesDocumentation ability to write clear runbooks, incident reports, and post-mortems
More at IH