Source description
About the role
Role Summary We are seeking an experienced technology leader to head the operations of a large-scale AI Cloud and High-Performance Computing (HPC) infrastructure. The role will be responsible for ensuring the availability, reliability, scalability, and operational excellence of mission-critical GPU, cloud, storage, and networking platforms supporting enterprise AI workloads. Key Responsibilities Lead 24x7 operations of AI cloud, GPU infrastructure, and HPC environments.Ensure high availability, performance, capacity planning, and operational excellence across compute, storage, networking, and cloud platforms.Drive Incident, Problem, Change, and Availability Management in line with ITIL best practices.Lead infrastructure lifecycle management, including upgrades, patching, capacity expansion, and technology refresh.Build and lead high-performing Cloud Operations, Infrastructure Operations, and NOC teams.Manage strategic relationships with OEMs, technology partners, and data center service providers.Drive automation, observability, operational governance, and continuous service improvement. Candidate Profile 18+ years of experience in Cloud Infrastructure, Data Center Operations, AI Infrastructure, or HPC environments.Proven experience managing large-scale enterprise or hyperscale infrastructure operations.Strong expertise in GPU infrastructure, Kubernetes/OpenShift, Enterprise Linux, high-performance networking, enterprise storage, cloud platforms, and ITIL-based service management.Experience leading large operations teams, managing vendors, and driving operational transformation. Preferred Skills AI InfrastructureNVIDIA GPU PlatformsHigh-Performance Computing (HPC)Cloud OperationsKubernetes / OpenShiftInfiniBand NetworkingEnterprise StorageITIL Service ManagementInfrastructure AutomationCapacity PlanningVendor ManagementIncident & Problem Management Education: Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline. Role Summary We are seeking an experienced technology leader to head the operations of a large-scale AI Cloud and High-Performance Computing (HPC) infrastructure. The role will be responsible for ensuring the availability, reliability, scalability, and operational excellence of mission-critical GPU, cloud, storage, and networking platforms supporting enterprise AI workloads. Key Responsibilities Lead 24x7 operations of AI cloud, GPU infrastructure, and HPC environments.Ensure high availability, performance, capacity planning, and operational excellence across compute, storage, networking, and cloud platforms.Drive Incident, Problem, Change, and Availability Management in line with ITIL best practices.Lead infrastructure lifecycle management, including upgrades, patching, capacity expansion, and technology refresh.Build and lead high-performing Cloud Operations, Infrastructure Operations, and NOC teams.Manage strategic relationships with OEMs, technology partners, and data center service providers.Drive automation, observability, operational governance, and continuous service improvement. Candidate Profile 18+ years of experience in Cloud Infrastructure, Data Center Operations, AI Infrastructure, or HPC environments.Proven experience managing large-scale enterprise or hyperscale infrastructure operations.Strong expertise in GPU infrastructure, Kubernetes/OpenShift, Enterprise Linux, high-performance networking, enterprise storage, cloud platforms, and ITIL-based service management.Experience leading large operations teams, managing vendors, and driving operational transformation. Preferred Skills AI InfrastructureNVIDIA GPU PlatformsHigh-Performance Computing (HPC)Cloud OperationsKubernetes / OpenShiftInfiniBand NetworkingEnterprise StorageITIL Service ManagementInfrastructure AutomationCapacity PlanningVendor ManagementIncident & Problem Management Education: Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.
More at LinkCxO (The CxO's Marketplace)