Padmi

Site Reliability Engineer

IndiaPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at Respironics Inc

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: As a Site Reliability Engineer (SRE) at our company, your primary responsibility will be to design and scale observability frameworks across cloud environments. You will define and manage SLIs/SLOs to ensure high availability, performance, and reliability. Additionally, you will be expected to build proactive, AI-driven monitoring systems to detect anomalies and predict failures. Your role will involve developing automation and self-healing capabilities to reduce manual intervention and improve system resilience. Collaboration with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency will be a key aspect of your responsibilities. Continuous improvement through incident analysis, performance tuning, and reliability enhancements will also be part of your daily tasks. Key Responsibilities: - Design and scale observability frameworks (metrics, logs, traces, event streams) across cloud environments - Define and manage SLIs/SLOs for high availability, performance, and reliability - Build proactive, AI-driven monitoring systems to detect anomalies and predict failures - Develop automation and self-healing capabilities to improve system resilience - Enable event-driven operations by integrating with tools like Service Now, Pager Duty, and Slack - Collaborate with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency - Drive continuous improvement through incident analysis, performance tuning, and reliability enhancements Qualifications Required: - Minimum 8+ years of experience in SRE/Cloud/Platform Engineering with AWS production environment experience - Expertise in Prometheus, Grafana, Datadog, Open Telemetry, Cloud Watch, and managing SLIs/SLOs - Strong skills in Python, Go, or Bash for building automation and self-healing systems - Experience with distributed systems, microservices, Docker, and Kubernetes - Knowledge of event-driven operations, incident tools (Service Now, Pager Duty, Slack), and root cause analysis - Experience working with cross-functional teams to drive performance, security, and cost optimization (Fin Ops) About the Company: Our company is a health technology company that believes in providing quality healthcare to everyone. We value teamwork and collaboration to make a positive impact on people's lives. If you are passionate about making a difference and have relevant experience, we encourage you to apply for this role. Role Overview: As a Site Reliability Engineer (SRE) at our company, your primary responsibility will be to design and scale observability frameworks across cloud environments. You will define and manage SLIs/SLOs to ensure high availability, performance, and reliability. Additionally, you will be expected to build proactive, AI-driven monitoring systems to detect anomalies and predict failures. Your role will involve developing automation and self-healing capabilities to reduce manual intervention and improve system resilience. Collaboration with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency will be a key aspect of your responsibilities. Continuous improvement through incident analysis, performance tuning, and reliability enhancements will also be part of your daily tasks. Key Responsibilities: - Design and scale observability frameworks (metrics, logs, traces, event streams) across cloud environments - Define and manage SLIs/SLOs for high availability, performance, and reliability - Build proactive, AI-driven monitoring systems to detect anomalies and predict failures - Develop automation and self-healing capabilities to improve system resilience - Enable event-driven operations by integrating with tools like Service Now, Pager Duty, and Slack - Collaborate with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency - Drive continuous improvement through incident analysis, performance tuning, and reliability enhancements Qualifications Required: - Minimum 8+ years of experience in SRE/Cloud/Platform Engineering with AWS production environment experience - Expertise in Prometheus, Grafana, Datadog, Open Telemetry, Cloud Watch, and managing SLIs/SLOs - Strong skills in Python, Go, or Bash for building automation and self-healing systems - Experience with distributed systems, microservices, Docker, and Kubernetes - Knowledge of event-driven operations, incident tools (Service Now, Pager Duty, Slack), and root cause analysis - Experience working with cross-functional teams to drive performance, security, and cost optimization (Fin Ops) About the Company: Our company is a health technology company that believes in providing quality healthcare to everyone. We value teamwork and collaboration to make a positive impact on people's lives. If you are passionate about making a difference and have relevant experience, we encourage you to apply for this role.

One address, no account. We’ll tell you when matching roles go live.

More at Respironics Inc

Related open roles

View all roles