Source description
About the role
About the Role We are seeking a highly experienced Sr DevOps Engineer Compute Platforms with strong operational expertise in enterprise compute platform s to implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud environment. This role focuses on the deployment, automation, support, and continuous improvement of large-scale compute environments spanning bare metal infrastructure, virtualization, private cloud, and Kubernetes platforms using Infrastructure-as-Code and GitOps practices. This is a deeply technical role requiring expert-level understanding of compute hardware management, Kubernetes, OpenStack, hypervisors and extensive working knowledge on Linux Operating systems. You will also collaborate with platform and SRE teams to maintain secure, performant, and multi-tenant-isolated services that serve high-throughput, mission-critical applications. Key Responsibilities Operate and support enterprise compute platforms across hardware, OS, virtualization, and container orchestration layers Develop and improve automation for provisioning, patching, upgrades, validation, and routine platform work using tools such as Ansible, Terraform, Helm, Git, Bash, Python, or similar technologies Deploy and Support Ubuntu-based systems, hypervisors, Kubernetes clusters, and related platform services in day-to-day operations Implement and maintain PXE-based provisioning environments leveraging Redfish APIs for large-scale server deployments Operate and support virtualization and private cloud platforms, including KVM on Ubuntu, OpenStack environments and Harvester HCI Monitor system performance, capacity, and availability; proactively address reliability risks Troubleshoot complex cross-stack issues spanning hardware, OS, virtualization, OpenStack, and Kubernetes Manage to SLAs, KPIs and error budgets Participate in on-call escalation support for complex platform-related issues Collaborate globally on change management, documentation, and operational best practices. Develop and maintain runbooks, operational procedures, and technical documentation Minimum Qualifications 6 years of experience as a DevOps Engineer, Site Reliability Engineer, or Infrastructure Operations Engineer with a strong focus on compute Strong hands-on experience operating bare metal compute environments at scale Experience with PXE boot, automated OS provisioning, and server imaging systems Practical experience supporting Bare Metal as a Service (BMaaS) platforms leveraging Redfish APIs Strong Linux administration skills, especially with Ubuntu Experience building and supporting CI/CD pipelines for infrastructure and platform automation Proficiency with Infrastructure as Code tools (e.g., Terraform, Ansible, or similar) Operational experience with virtualization and private cloud platforms, including KVM on Ubuntu, OpenStack operations and troubleshooting, Harvester HCI Experience deploying and operating production Kubernetes environments Strong scripting skills in Python, Bash, or similar languages Strong understanding on SRE functions like toil reduction, error budgets and meeting SLAs Proven troubleshooting and root cause analysis skills in complex distributed systems Excellent written and verbal communication skills Bachelor s degree in computer science or equivalent professional experience. Preferred Qualifications OpenStack, Ubuntu KVM administration. BareMetal as a Service (PXE, Redfish) Kubernetes on BareMetal Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at Five9