Source description
About the role
Role Overview: As a Site Reliability Engineer (SRE) at our company, your primary responsibility will be to design and scale observability frameworks across cloud environments. You will define and manage SLIs/SLOs to ensure high availability, performance, and reliability. Additionally, you will be expected to build proactive, AI-driven monitoring systems to detect anomalies and predict failures. Your role will involve developing automation and self-healing capabilities to reduce manual intervention and improve system resilience. Collaboration with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency will be a key aspect of your responsibilities. Continuous improvement through incident analysis, performance tuning, and reliability enhancements will also be part of your daily tasks. Key Responsibilities: - Design and scale observability frameworks (metrics, logs, traces, event streams) across cloud environments - Define and manage SLIs/SLOs for high availability, performance, and reliability - Build proactive, AI-driven monitoring systems to detect anomalies and predict failures - Develop automation and self-healing capabilities to improve system resilience - Enable event-driven operations by integrating with tools like Service Now, Pager Duty, and Slack - Collaborate with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency - Drive continuous improvement through incident analysis, performance tuning, and reliability enhancements Qualifications Required: - Minimum 8+ years of experience in SRE/Cloud/Platform Engineering with AWS production environment experience - Expertise in Prometheus, Grafana, Datadog, Open Telemetry, Cloud Watch, and managing SLIs/SLOs - Strong skills in Python, Go, or Bash for building automation and self-healing systems - Experience with distributed systems, microservices, Docker, and Kubernetes - Knowledge of event-driven operations, incident tools (Service Now, Pager Duty, Slack), and root cause analysis - Experience working with cross-functional teams to drive performance, security, and cost optimization (Fin Ops) About the Company: Our company is a health technology company that believes in providing quality healthcare to everyone. We value teamwork and collaboration to make a positive impact on people's lives. If you are passionate about making a difference and have relevant experience, we encourage you to apply for this role. Role Overview: As a Site Reliability Engineer (SRE) at our company, your primary responsibility will be to design and scale observability frameworks across cloud environments. You will define and manage SLIs/SLOs to ensure high availability, performance, and reliability. Additionally, you will be expected to build proactive, AI-driven monitoring systems to detect anomalies and predict failures. Your role will involve developing automation and self-healing capabilities to reduce manual intervention and improve system resilience. Collaboration with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency will be a key aspect of your responsibilities. Continuous improvement through incident analysis, performance tuning, and reliability enhancements will also be part of your daily tasks. Key Responsibilities: - Design and scale observability frameworks (metrics, logs, traces, event streams) across cloud environments - Define and manage SLIs/SLOs for high availability, performance, and reliability - Build proactive, AI-driven monitoring systems to detect anomalies and predict failures - Develop automation and self-healing capabilities to improve system resilience - Enable event-driven operations by integrating with tools like Service Now, Pager Duty, and Slack - Collaborate with engineering, Sec Ops, and Fin Ops teams to enhance reliability, security, and cost efficiency - Drive continuous improvement through incident analysis, performance tuning, and reliability enhancements Qualifications Required: - Minimum 8+ years of experience in SRE/Cloud/Platform Engineering with AWS production environment experience - Expertise in Prometheus, Grafana, Datadog, Open Telemetry, Cloud Watch, and managing SLIs/SLOs - Strong skills in Python, Go, or Bash for building automation and self-healing systems - Experience with distributed systems, microservices, Docker, and Kubernetes - Knowledge of event-driven operations, incident tools (Service Now, Pager Duty, Slack), and root cause analysis - Experience working with cross-functional teams to drive performance, security, and cost optimization (Fin Ops) About the Company: Our company is a health technology company that believes in providing quality healthcare to everyone. We value teamwork and collaboration to make a positive impact on people's lives. If you are passionate about making a difference and have relevant experience, we encourage you to apply for this role.
More at Respironics Inc