Padmi

Site Reliability Engineer

IndiaPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at IH

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: As an experienced Site Reliability Engineer (SRE) with over 3 years of experience, you will be responsible for incident response and supporting production services at scale on a major cloud provider such as Azure, AWS, or GCP. Your role may also involve working with API gateway or proxy infrastructure, providing you with an opportunity to enhance your skills further. Key Responsibilities: - Write PromQL queries from scratch, build dashboards, and interpret results using Prometheus & Grafana - Communicate incidents in a structured, calm, and clear manner both in written and verbal form, especially under pressure - Understand observability fundamentals including metrics, logs, and traces, and know when to use each - Demonstrate cloud fluency by having a working knowledge of at least one major cloud provider's managed services and common failure modes - Document effectively by writing clear runbooks, incident reports, and post-mortems Qualifications Required: - Minimum of 3 years of experience in an SRE role - Proven track record in incident response - Experience supporting production services at scale on a major cloud provider (Azure, AWS, or GCP) - Experience working with API gateway or proxy infrastructure is a plus Additional Company Details: N/A Role Overview: As an experienced Site Reliability Engineer (SRE) with over 3 years of experience, you will be responsible for incident response and supporting production services at scale on a major cloud provider such as Azure, AWS, or GCP. Your role may also involve working with API gateway or proxy infrastructure, providing you with an opportunity to enhance your skills further. Key Responsibilities: - Write PromQL queries from scratch, build dashboards, and interpret results using Prometheus & Grafana - Communicate incidents in a structured, calm, and clear manner both in written and verbal form, especially under pressure - Understand observability fundamentals including metrics, logs, and traces, and know when to use each - Demonstrate cloud fluency by having a working knowledge of at least one major cloud provider's managed services and common failure modes - Document effectively by writing clear runbooks, incident reports, and post-mortems Qualifications Required: - Minimum of 3 years of experience in an SRE role - Proven track record in incident response - Experience supporting production services at scale on a major cloud provider (Azure, AWS, or GCP) - Experience working with API gateway or proxy infrastructure is a plus Additional Company Details: N/A

One address, no account. We’ll tell you when matching roles go live.

More at IH

Related open roles

View all roles