Source description
About the role
As a Senior Infrastructure Test & Validation Engineer at Our Client Corporation, you will play a crucial role in leading the Zero-Touch Validation, Upgrade, and Certification automation of the on-prem GPU cloud platform. Your primary focus will be on ensuring stability, performance, and conformance across the entire stack using automated GitOps-based validation pipelines. Your deep expertise in infrastructure engineering and hands-on skills in tools like Sonobuoy, LitmusChaos, k6, and pytest will be instrumental in the success of this role. Key Responsibilities: - Design and implement automated, GitOps-compliant pipelines for validation and certification of the GPU cloud stack. - Integrate Sonobuoy for Kubernetes conformance and certification testing. - Orchestrate chaos engineering workflows using LitmusChaos to validate system resilience. - Implement performance testing suites using k6 and system-level benchmarks integrated into CI/CD pipelines. - Develop end-to-end test frameworks using pytest and/or Go, focusing on cluster lifecycle events, upgrade paths, and GPU workloads. - Ensure test coverage across multiple dimensions including conformance, performance, fault injection, and post-upgrade validation. - Build and maintain dashboards and reporting for automated test results. - Collaborate with infrastructure, SRE, and platform teams to embed testing and validation early in the deployment lifecycle. - Own quality assurance gates for all automation-driven deployments. Required Skills & Experience: - 10+ years of hands-on experience in infrastructure engineering, systems validation, or SRE roles. - Proficiency in pytest, Go, k6 scripting, and automation frameworks integration (Sonobuoy, LitmusChaos). - Strong understanding of Kubernetes architecture, upgrade patterns, and operational risks. - Experience with GitOps workflows and CI/CD systems for test orchestration. - Solid scripting and automation experience in Python, Bash, or Go. - Familiarity with GPU-based infrastructure and its performance characteristics is a strong plus. - Strong debugging, root cause analysis, and incident investigation skills. Join Our Client Corporation to be a part of a dynamic team that prioritizes execution early in the process for impactful results. Your expertise will contribute to creatively solving pressing business challenges and moving clients to the forefront of their industry. As a Senior Infrastructure Test & Validation Engineer at Our Client Corporation, you will play a crucial role in leading the Zero-Touch Validation, Upgrade, and Certification automation of the on-prem GPU cloud platform. Your primary focus will be on ensuring stability, performance, and conformance across the entire stack using automated GitOps-based validation pipelines. Your deep expertise in infrastructure engineering and hands-on skills in tools like Sonobuoy, LitmusChaos, k6, and pytest will be instrumental in the success of this role. Key Responsibilities: - Design and implement automated, GitOps-compliant pipelines for validation and certification of the GPU cloud stack. - Integrate Sonobuoy for Kubernetes conformance and certification testing. - Orchestrate chaos engineering workflows using LitmusChaos to validate system resilience. - Implement performance testing suites using k6 and system-level benchmarks integrated into CI/CD pipelines. - Develop end-to-end test frameworks using pytest and/or Go, focusing on cluster lifecycle events, upgrade paths, and GPU workloads. - Ensure test coverage across multiple dimensions including conformance, performance, fault injection, and post-upgrade validation. - Build and maintain dashboards and reporting for automated test results. - Collaborate with infrastructure, SRE, and platform teams to embed testing and validation early in the deployment lifecycle. - Own quality assurance gates for all automation-driven deployments. Required Skills & Experience: - 10+ years of hands-on experience in infrastructure engineering, systems validation, or SRE roles. - Proficiency in pytest, Go, k6 scripting, and automation frameworks integration (Sonobuoy, LitmusChaos). - Strong understanding of Kubernetes architecture, upgrade patterns, and operational risks. - Experience with GitOps workflows and CI/CD systems for test orchestration. - Solid scripting and automation experience in Python, Bash, or Go. - Familiarity with GPU-based infrastructure and its performance characteristics is a strong plus. - Strong debugging, root cause analysis, and incident investigation skills. Join Our Client Corporation to be a part of a dynamic team that prioritizes execution early in the process for impactful results. Your expertise will contribute to creatively solving pressing business challenges and moving clients to the forefront of their industry.
More at People Prime Worldwide