Source description
About the role
Technologyor technology that breathes life into work SHL, People Science. People Answers. Are you a seasoned Site Reliability Engineer with a flair for innovation Are you ready to shape the future of talent assessment and empower organizations to unlock their full potential If so, we want you to be a part of the SHL Team! As a Site Reliability Engineer, Were seeking an experienced Site Reliability Engineer (SRE) to help keep our systems fast, reliable, and . This role is ideal for someone about driving uptime, performance, and automation while fostering a culture of continuous improvement An excellent benefit package is offered in a culture where career development, with ongoing manager guidance, collaboration, flexibility, diversity, and inclusivity are all intrinsic to our culture. There is a huge investment in SHL currently so theres no better time to become a part of something transformational. What you will be doing: Apply SRE conceptsavailability, latency, error rates, and saturationto maintain reliable, scalable systems Champion a zero-downtime mindset, ensuring high availability with minimal service disruptions Define and guide SLIs, SLOs, and SLAs, establishing meaningful error budgets and tracking them diligently Enhance API performance through detailed evaluation of latency and percentile metrics (p50, p90, p95, p99) Handle on-call rotations and production assist for large-scale distributed systems Monitoring & Observability Work hands-on with observability tools such as Prometheus, Grafana, Datadog, New Relic, or the ELK Stack Implement metrics, logging, and distributed tracing for complete system visibility Design custom dashboards, alerts, and proactive monitoring workflows Use APM tools and real-time performance monitoring for early detection and prevention of issues What we are looking for from you: Essential: Proven skills in capacity planning, performance tuning, and identifying system bottlenecks Scripting expertise in Python or Go to automate repetitive operational tasks Experience conducting post-incident reviews and leading blameless postmortems to drive process improvements Desirable: Strong experience in incident management with a focus on reducing MTTR and preventing recurrences Get in touch: Find out how this one-off opportunity can help you to achieve your career goals by making an application to our knowledgeable and friendly Talent Acquisition team. Choose a new path with SHL. Technologyor technology that breathes life into work SHL, People Science. People Answers. Are you a seasoned Site Reliability Engineer with a flair for innovation Are you ready to shape the future of talent assessment and empower organizations to unlock their full potential If so, we want you to be a part of the SHL Team! As a Site Reliability Engineer, Were seeking an experienced Site Reliability Engineer (SRE) to help keep our systems fast, reliable, and . This role is ideal for someone about driving uptime, performance, and automation while fostering a culture of continuous improvement An excellent benefit package is offered in a culture where career development, with ongoing manager guidance, collaboration, flexibility, diversity, and inclusivity are all intrinsic to our culture. There is a huge investment in SHL currently so theres no better time to become a part of something transformational. What you will be doing: Apply SRE conceptsavailability, latency, error rates, and saturationto maintain reliable, scalable systems Champion a zero-downtime mindset, ensuring high availability with minimal service disruptions Define and guide SLIs, SLOs, and SLAs, establishing meaningful error budgets and tracking them diligently Enhance API performance through detailed evaluation of latency and percentile metrics (p50, p90, p95, p99) Handle on-call rotations and production assist for large-scale distributed systems Monitoring & Observability Work hands-on with observability tools such as Prometheus, Grafana, Datadog, New Relic, or the ELK Stack Implement metrics, logging, and distributed tracing for complete system visibility Design custom dashboards, alerts, and proactive monitoring workflows Use APM tools and real-time performance monitoring for early detection and prevention of issues What we are looking for from you: Essential: Proven skills in capacity planning, performance tuning, and identifying system bottlenecks Scripting expertise in Python or Go to automate repetitive operational tasks Experience conducting post-incident reviews and leading blameless postmortems to drive process improvements Desirable: Strong experience in incident management with a focus on reducing MTTR and preventing recurrences Get in touch: Find out how this one-off opportunity can help you to achieve your career goals by making an application to our knowledgeable and friendly Talent Acquisition team. Choose a new path with SHL.
More at SHL