Source description
About the role
We are seeking a highly skilled Site Reliability Engineer (SRE) to join our dynamic infrastructure and platform team. The ideal candidate will have strong expertise in Kubernetes, Jenkins, Docker, and scripting (Python/Bash) , with a passion for building resilient, scalable, and automated infrastructure systems. As an SRE, you will play a key role in ensuring system uptime, performance, and efficiency while collaborating closely with development and operations teams. Roles & Responsibilities: Design, deploy, and maintain highly available and scalable Kubernetes clusters Develop and manage CI/CD pipelines using Jenkins Create, manage, and monitor containerized applications using Docker Write and maintain automation scripts in Python and Bash to eliminate manual tasks Implement monitoring and alerting using Prometheus and Alertmanager Manage infrastructure as code with tools like Terraform Support GitOps workflows using ArgoCD for continuous deployment Administer secrets and sensitive data using Vault Maintain and optimize Openshift (OCP) environments Participate in incident response, postmortems , and ensure system reliability through SLAs/SLOs Drive reliability and performance improvements across systems and services Primary Skills: Site Reliability Engineering (SRE) best practices Kubernetes (deployment, scaling, configuration) Jenkins (CI/CD pipeline management) Docker (image creation, orchestration) Scripting: Python and Bash Secondary / Good-to-Have Skills: Terraform (Infrastructure as Code) OpenShift (OCP) ArgoCD (GitOps) Vault (Secrets management) Prometheus & Alertmanager (Monitoring and alerting) Preferred Qualifications: Bachelor's degree in Computer Science, Engineering, or related field Certifications in Kubernetes, Terraform, or Cloud DevOps (AWS/GCP/Azure) Experience working in DevOps, SRE, or Platform Engineering roles in large-scale environments
More at Peoplefy Infosolutions