Source description
About the role
Role Overview: As a Site Reliability Engineer (SRE) - Technical Leader at Cisco Meraki, you will play a crucial role in designing, operating, and scaling Kubernetes-based platforms supporting various environments. You will work at the intersection of software engineering and infrastructure, collaborating closely with engineers to ensure the platform is resilient, observable, compliant, and developer-friendly. Key Responsibilities: - Design, build, and operate production-grade Kubernetes platforms in regulated and non-regulated environments. - Improve system reliability through automation, thoughtful design, and continuous iteration. - Define and drive SLOs, SLIs, and error budgets to guide reliability decisions. - Build and evolve secure, scalable, and easy-to-use CI/CD pipelines. - Implement robust observability (metrics, logs, traces) to enhance system understandability. - Reduce operational toil by automating repetitive processes and improving workflows. - Collaborate with security and compliance teams to meet Compliance requirements without hindering developer velocity. - Support audit processes, including documentation, controls implementation, and audit readiness. - Participate in on-call rotations supporting customer requests and paging alerts. - Engage in incident response, blameless postmortems, and continuous improvement efforts. - Contribute to shaping a platform that engineers find enjoyable to use. Qualifications Required: - 10+ years of experience in SRE, DevOps, or infrastructure engineering. - Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream). - Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals. - Proficiency in Infrastructure as Code, with Terraform preferred. - Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD). - Proficiency in scripting or programming languages such as Python or Go. - Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK). - Working knowledge of compliance frameworks (e.g., PCI, ISO). Additional Company Details: Cisco is revolutionizing how data and infrastructure connect and protect organizations in the AI era and beyond. With a history of innovation spanning 40 years, Cisco creates solutions that empower the collaboration between humans and technology. The company's solutions offer unparalleled security, visibility, and insights across the digital footprint. Cisco fosters a culture of experimentation and collaboration, providing endless opportunities for growth and development on a global scale. Please note: This job posting has been aggregated from an external source, and role details are subject to change. It is recommended to verify the latest information on the company website before applying. Role Overview: As a Site Reliability Engineer (SRE) - Technical Leader at Cisco Meraki, you will play a crucial role in designing, operating, and scaling Kubernetes-based platforms supporting various environments. You will work at the intersection of software engineering and infrastructure, collaborating closely with engineers to ensure the platform is resilient, observable, compliant, and developer-friendly. Key Responsibilities: - Design, build, and operate production-grade Kubernetes platforms in regulated and non-regulated environments. - Improve system reliability through automation, thoughtful design, and continuous iteration. - Define and drive SLOs, SLIs, and error budgets to guide reliability decisions. - Build and evolve secure, scalable, and easy-to-use CI/CD pipelines. - Implement robust observability (metrics, logs, traces) to enhance system understandability. - Reduce operational toil by automating repetitive processes and improving workflows. - Collaborate with security and compliance teams to meet Compliance requirements without hindering developer velocity. - Support audit processes, including documentation, controls implementation, and audit readiness. - Participate in on-call rotations supporting customer requests and paging alerts. - Engage in incident response, blameless postmortems, and continuous improvement efforts. - Contribute to shaping a platform that engineers find enjoyable to use. Qualifications Required: - 10+ years of experience in SRE, DevOps, or infrastructure engineering. - Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream). - Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals. - Proficiency in Infrastructure as Code, with Terraform preferred. - Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD). - Proficiency in scripting or programming languages such as Python or Go. - Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK). - Working knowledge of compliance frameworks (e.g., PCI, ISO). Additional Company Details: Cisco is revolutionizing how data and infrastructure
More at Cisco