Source description
About the role
As a Site Reliability Engineer (SRE) at our company, you will be responsible for the daily operations, architectural resilience, and implementation of SRE principles in a large-scale environment. Your role will involve understanding various technology domains and their interactions to achieve business objectives. You will play a crucial part in enhancing the reliability, performance, and efficiency of our Applications and Services. Key Responsibilities: - Foster a culture of transparency, innovation, and accountability to encourage continuous improvement. - Communicate progress and impact of SRE initiatives to stakeholders at all levels. - Ensure compliance with all relevant requirements in a highly regulated environment. - Oversee advanced recovery testing, including various tests like Production Swing Tests, Data Recovery Tests, and chaos engineering practices. - Drive the adoption and development of automation solutions to minimize recovery time. - Collaborate with development teams to leverage cloud native services and enhance application reliability. - Collaborate across the organization to develop and scale observability solutions using modern tools for metrics, logging, and tracing. - Partner with development teams to instrument applications effectively, providing insights into system health and performance. Essential Skills: - 13+ years of deep understanding of SRE concepts such as SLOs, SLIs, error budgets, and toil reduction. - Experience with Disaster Recovery planning, resiliency testing, and fault-tolerant distributed system design. - Proficiency in deploying, managing, and troubleshooting applications on OpenShift/Kubernetes. - Hands-on experience with modern observability tools like Prometheus, Grafana, Loki, Mimir, Tempo, and AppDynamics. - Experience with Infrastructure as Code (IaC), configuration management, and automation tools like Ansible and Terraform. - Experience in creating, modifying, and managing Helm charts for application deployment. Desired Skills: - Experience with major public cloud providers like Google Cloud, AWS, Azure. - Proven experience in delivering software and infrastructure using Agile frameworks. - Experience presenting technical strategy to senior and executive-level audiences. - Experience in writing or maintaining code in Java, Python, Gco, or similar languages. Qualifications: - Significant professional experience in production management, software development, or a related field with a strong focus on Site Reliability Engineering. - Expertise in analyzing complex application, database, network, and OS issues within large-scale systems. - A service-oriented attitude with excellent problem-solving and strategic thinking skills. - Strong communication and diplomacy skills, with the ability to work effectively across multiple business and technical teams. Citi is an equal opportunity employer, and all qualified candidates will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law. As a Site Reliability Engineer (SRE) at our company, you will be responsible for the daily operations, architectural resilience, and implementation of SRE principles in a large-scale environment. Your role will involve understanding various technology domains and their interactions to achieve business objectives. You will play a crucial part in enhancing the reliability, performance, and efficiency of our Applications and Services. Key Responsibilities: - Foster a culture of transparency, innovation, and accountability to encourage continuous improvement. - Communicate progress and impact of SRE initiatives to stakeholders at all levels. - Ensure compliance with all relevant requirements in a highly regulated environment. - Oversee advanced recovery testing, including various tests like Production Swing Tests, Data Recovery Tests, and chaos engineering practices. - Drive the adoption and development of automation solutions to minimize recovery time. - Collaborate with development teams to leverage cloud native services and enhance application reliability. - Collaborate across the organization to develop and scale observability solutions using modern tools for metrics, logging, and tracing. - Partner with development teams to instrument applications effectively, providing insights into system health and performance. Essential Skills: - 13+ years of deep understanding of SRE concepts such as SLOs, SLIs, error budgets, and toil reduction. - Experience with Disaster Recovery planning, resiliency testing, and fault-tolerant distributed system design. - Proficiency in deploying, managing, and troubleshooting applications on OpenShift/Kubernetes. - Hands-on experience with modern observability tools like Prometheus, Grafana, Loki, Mimir, Tempo, and AppDynamics. - Experience with Infrastructure as Code (I
More at Professional