Source description
About the role
Role Overview: As a Senior Site Reliability Engineer (SRE) at Cirium, you will have the opportunity to work on impactful projects that enhance reliability and reduce manual work through automation. Your role will involve leveraging your experience across various SRE practices to maintain resilient, distributed systems and automate processes to protect critical services. Additionally, you will participate in on-call rotations, offer guidance and support to colleagues, and contribute to shaping an inclusive culture of technical excellence and continuous learning. Key Responsibilities: - Lead efforts to automate manual and repetitive tasks, contributing to resilient and reliable systems. - Develop and implement self-healing infrastructure solutions to enhance operational efficiency and reduce incidents. - Create and maintain automation and tools to promote system performance and uptime. - Support post-release validation and operational readiness for new deployments. - Provide occasional support outside of standard hours as needed for major releases or critical changes, with consideration for work-life balance. - Design infrastructure following best practices for scalability, fault tolerance, and security. - Define and manage Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to ensure reliable services. - Collaborate with engineering teams to enhance deployment pipelines and make recommendations for improved architecture, release speed, and productivity. Qualifications Required: - Professional experience in a Site Reliability Engineering (SRE), DevOps, or related technical role. - Familiarity with cloud platforms, especially AWS services such as EC2, ECS, Lambda, Redshift, PostgreSQL, and AWS OpenSearch. - Hands-on experience with Infrastructure as Code (IaC) tools like Terraform for automating and managing cloud infrastructure. - Experience with containerization technologies such as Docker; Kubernetes experience is advantageous. - Proficiency in CI/CD pipelines, build/release management tools, and deployment/operational understanding of applications like Apache Airflow, .NET applications, and enterprise systems. - Knowledge of configuration management tools like Puppet, Ansible, or equivalent systems. - Experience with monitoring, alerting, and observability tools such as Elasticsearch, Grafana, OpenTelemetry, GitHub Actions, Azure DevOps, TeamCity, Jenkins, or similar platforms. - Ability to work independently, collaborate effectively with cross-functional teams, and communicate technical issues clearly. - Openness to learning new technologies, exploring modern engineering practices, and driving innovation and continuous improvement initiatives. - Familiarity with AI-assisted developer tools like GitHub Copilot is a plus. - Relevant certifications in AWS, Kubernetes, or related areas are appreciated but not mandatory. (Note: Additional details about the company were not included in the provided job description.) Role Overview: As a Senior Site Reliability Engineer (SRE) at Cirium, you will have the opportunity to work on impactful projects that enhance reliability and reduce manual work through automation. Your role will involve leveraging your experience across various SRE practices to maintain resilient, distributed systems and automate processes to protect critical services. Additionally, you will participate in on-call rotations, offer guidance and support to colleagues, and contribute to shaping an inclusive culture of technical excellence and continuous learning. Key Responsibilities: - Lead efforts to automate manual and repetitive tasks, contributing to resilient and reliable systems. - Develop and implement self-healing infrastructure solutions to enhance operational efficiency and reduce incidents. - Create and maintain automation and tools to promote system performance and uptime. - Support post-release validation and operational readiness for new deployments. - Provide occasional support outside of standard hours as needed for major releases or critical changes, with consideration for work-life balance. - Design infrastructure following best practices for scalability, fault tolerance, and security. - Define and manage Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to ensure reliable services. - Collaborate with engineering teams to enhance deployment pipelines and make recommendations for improved architecture, release speed, and productivity. Qualifications Required: - Professional experience in a Site Reliability Engineering (SRE), DevOps, or related technical role. - Familiarity with cloud platforms, especially AWS services such as EC2, ECS, Lambda, Redshift, PostgreSQL, and AWS OpenSearch. - Hands-on experience with Infrastructure as Code (IaC) tools like Terraform for automating and managing cloud infrastructure. - Experience with containerization technologies such as Docker; Kubernetes experience is advantageous. - Proficiency in CI/CD pipelines, b
More at LexisNexis Risk Solutions