Padmi
NVIDIA logo
NVIDIA

GPU computing · AI infrastructure

Senior Platform Engineer, Network Infrastructure Hyderabad (India)

HyderabadPosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at NVIDIA

Opens the source posting on shine.com

Source description

About the role

View original

Cloud Foundations Reliability (CFR) is part of NVIDIAs Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIAs global network across data centers, colocation facilities, and cloud settings. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments. We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility. What Youll Be Doing: Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments. Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery. Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and protected multi-cluster delivery through GitOps. Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features. Diagnose complex .

One address, no account. We’ll tell you when matching roles go live.

More at NVIDIA

Related open roles

View all roles