Source description
About the role
About the Role: We are seeking a highly motivated and experienced Platform Reliability Engineer (PRE) to ensure the performance, reliability, and scalability of our core platform infrastructure. In this role, you will work at the intersection of software engineering and systems engineering to build resilient systems, automate operational processes, and drive platform efficiency. You will play a critical role in minimizing downtime, reducing manual work, and improving the developer and user experience across the platform. Key Responsibilities: - Design, build, and maintain reliable infrastructure and platforms that scale with business needs. - Monitor, measure, and improve system reliability and performance using observability tools. - Collaborate with development and DevOps teams to implement best practices in reliability engineering. - Build automation and self-healing systems to reduce manual operations and increase platform uptime. - Define and enforce service-level indicators (SLIs), objectives (SLOs), and error budgets. - Participate in capacity planning, incident response, and post-incident reviews. - Implement CI/CD pipelines to streamline deployment processes and reduce failure rates. - Enhance system observability through better logging, monitoring, and alerting practices. - Advocate for and implement improvements in infrastructure-as-code and platform scalability. Preferred Skills: - Experience working with microservices and distributed systems. - Knowledge of service mesh technologies (e.g., Istio, Linkerd). - Experience with databases (SQL and NoSQL) and tuning for high availability. - Certifications in cloud technologies (AWS, Azure, GCP) or Kubernetes. - Exposure to chaos engineering, game days, or failure injection strategies. About the Role: We are seeking a highly motivated and experienced Platform Reliability Engineer (PRE) to ensure the performance, reliability, and scalability of our core platform infrastructure. In this role, you will work at the intersection of software engineering and systems engineering to build resilient systems, automate operational processes, and drive platform efficiency. You will play a critical role in minimizing downtime, reducing manual work, and improving the developer and user experience across the platform. Key Responsibilities: - Design, build, and maintain reliable infrastructure and platforms that scale with business needs. - Monitor, measure, and improve system reliability and performance using observability tools. - Collaborate with development and DevOps teams to implement best practices in reliability engineering. - Build automation and self-healing systems to reduce manual operations and increase platform uptime. - Define and enforce service-level indicators (SLIs), objectives (SLOs), and error budgets. - Participate in capacity planning, incident response, and post-incident reviews. - Implement CI/CD pipelines to streamline deployment processes and reduce failure rates. - Enhance system observability through better logging, monitoring, and alerting practices. - Advocate for and implement improvements in infrastructure-as-code and platform scalability. Preferred Skills: - Experience working with microservices and distributed systems. - Knowledge of service mesh technologies (e.g., Istio, Linkerd). - Experience with databases (SQL and NoSQL) and tuning for high availability. - Certifications in cloud technologies (AWS, Azure, GCP) or Kubernetes. - Exposure to chaos engineering, game days, or failure injection strategies.
More at TI Steps