Padmi

TitleKubernetes + Terraform + Python

MumbaiPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at LTM

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Kubernetes Platform Engineer with 10+ years of infrastructure experience, you will be responsible for designing and implementing the Zero-Touch Build, Upgrade, and Certification pipeline for the on-premises GPU cloud platform. Your main focus will be on automating the Kubernetes layer and its dependencies such as GPU drivers, networking, and runtime using 100% GitOps workflows. You will collaborate with various teams to deliver a fully declarative, scalable, and reproducible infrastructure stack from hardware to Kubernetes and platform services. Key Responsibilities: - Architect and implement GitOps-driven Kubernetes cluster lifecycle automation utilizing tools like kubeadm, ClusterAPI, Helm, and Argo CD. - Develop and manage declarative infrastructure components for GPU stack deployment, container runtime configuration, and networking layers. - Lead automation efforts for zero-touch upgrades and certification pipelines for Kubernetes clusters and associated workloads. - Maintain Git-backed sources of truth for all platform configurations and integrations. - Standardize deployment practices across multi-cluster GPU environments to ensure scalability, repeatability, and compliance. - Drive observability, testing, and validation in the continuous delivery process, including cluster conformance and GPU health checks. - Collaborate with infrastructure, security, and SRE teams to ensure seamless handoffs between lower layers and the Kubernetes platform. - Mentor junior engineers and contribute to the platform automation roadmap. Qualification Required: - 10+ years of hands-on experience in infrastructure engineering, with a strong focus on Kubernetes-based environments. - Proficiency in Kubernetes API, Helm templating, Argo CD GitOps integration, Go/Python scripting, and Containerd. - Deep knowledge and hands-on experience with Kubernetes cluster management, Argo CD for GitOps-based delivery, Helm for application and cluster add-on packaging, and Container as a container runtime. - Experience deploying and operating the NVIDIA GPU Operator or equivalent in production environments. - Solid understanding of CNI plugin ecosystems, network policies, and multi-tenant networking in Kubernetes. - Strong GitOps mindset with experience managing infrastructure as code through Git-based workflows. - Proven ability to scale and manage multi-cluster, GPU-accelerated workloads with high availability and security. - Solid scripting and automation skills in Bash, Python, or Go. - Familiarity with Linux internals, systemd, and OS-level tuning for container workloads. Bonus: - Experience with custom controllers, operators, or Kubernetes API extensions. - Contributions to Kubernetes or CNCF projects. - Exposure to service meshes, ingress controllers, or workload identity providers. As a Senior Kubernetes Platform Engineer with 10+ years of infrastructure experience, you will be responsible for designing and implementing the Zero-Touch Build, Upgrade, and Certification pipeline for the on-premises GPU cloud platform. Your main focus will be on automating the Kubernetes layer and its dependencies such as GPU drivers, networking, and runtime using 100% GitOps workflows. You will collaborate with various teams to deliver a fully declarative, scalable, and reproducible infrastructure stack from hardware to Kubernetes and platform services. Key Responsibilities: - Architect and implement GitOps-driven Kubernetes cluster lifecycle automation utilizing tools like kubeadm, ClusterAPI, Helm, and Argo CD. - Develop and manage declarative infrastructure components for GPU stack deployment, container runtime configuration, and networking layers. - Lead automation efforts for zero-touch upgrades and certification pipelines for Kubernetes clusters and associated workloads. - Maintain Git-backed sources of truth for all platform configurations and integrations. - Standardize deployment practices across multi-cluster GPU environments to ensure scalability, repeatability, and compliance. - Drive observability, testing, and validation in the continuous delivery process, including cluster conformance and GPU health checks. - Collaborate with infrastructure, security, and SRE teams to ensure seamless handoffs between lower layers and the Kubernetes platform. - Mentor junior engineers and contribute to the platform automation roadmap. Qualification Required: - 10+ years of hands-on experience in infrastructure engineering, with a strong focus on Kubernetes-based environments. - Proficiency in Kubernetes API, Helm templating, Argo CD GitOps integration, Go/Python scripting, and Containerd. - Deep knowledge and hands-on experience with Kubernetes cluster management, Argo CD for GitOps-based delivery, Helm for application and cluster add-on packaging, and Container as a container runtime. - Experience deploying and operating the NVIDIA GPU Operator or equivalent in production environments. - Solid understanding of CNI plugin ecosystems, n

One address, no account. We’ll tell you when matching roles go live.

More at LTM

Related open roles

View all roles