Padmi

Site Reliability Engineer II Performance - Chaos

HyderabadPosted 1 month ago
Software engineeringMid-levelFull Time; Regular
Apply at Quest Diagnostics

Opens the source posting on shine.com

Source description

About the role

View original

As a Site Reliability chaos Engineer, your role is to provide reliability engineering services through chaos engineering, and performance engineering techniques. Using monitoring, fault-injection, and performance tools, you will deliver detailed feedback to product owners and development teams. You will collaborate with cross-functional teams to design, build, automate, and maintain scalable, resilient infrastructure. Your responsibilities will include ensuring high availability, monitoring system performance, executing chaos experiments, and aiding support staff with resolving incidents. This role requires a strong background in scripting, cloud platforms, and a passion for optimizing operational efficiency and system resilience. You will use Site Reliability Engineering chaos practices to deliver a seamless user experience. Responsibilities: Design and develop custom, automated fault-injection playbooks to test system recovery mechanisms under pressure. Execute end-to-end chaos experiments (such as network latency, container crashes, and region failovers) across non-production environments. Analyze infrastructure bottlenecks and application dependencies exposed during simulated system failures. Draft comprehensive post-mortem reports and present findings to software development teams to guide code-hardening efforts. Construct custom dashboards in Grafana and Prometheus to visualize real-time system degradation during chaos testing. Implement automated chaos tests directly into active CI/CD deployment pipelines. Perform extensive performance and load testing using JMeter to baseline system stability before injecting faults. Configure advanced Dynatrace alerting policies to ensure simulated outages are immediately and accurately detected by monitoring tools. Simulate full Disaster Recovery (DR) and business continuity scenarios to verify the reliability of backup and failover procedures. Partner with software engineers to refactor application code for better fault tolerance and self-healing capabilities. Collaborate with the broader operations team to define system boundaries and minimize risk when executing testing windows. Build clean, modular automation scripts in Python, Shell, or JavaScript to provision test environments and schedule recurring chaos runs. .

One address, no account. We’ll tell you when matching roles go live.

More at Quest Diagnostics

Related open roles

View all roles