Source description
About the role
We are looking for an outstanding DevOps and Site Reliability Engineer to join the NVIDIA e-commerce team. You will be a key architect of our e-commerce platform, ensuring that our systems are scalable, resilient, and automated. The ideal candidate is a Terraform expert who views infrastructure as code (IaC) not just as a tool, but as a philosophy. You will bridge the gap between development and operations, focusing on system reliability, high availability, and the performance of our global e-commerce platform. What You'll Be Doing Architect and refine automated deployment Jenkins pipelines to ensure seamless, zero-downtime releases. Design, build, and maintain enterprise-scale infrastructure using Terraform. Establish modular, reusable patterns for AWS resources. Optimize and manage sophisticated AWS environments with a focus on cost-efficiency and security. Transition our monitoring from reactive to proactive using AI-powered observability tools (e.g., Datadog Watchdog) for automated root cause analysis (RCA) and anomaly detection. Define and monitor SLOs and SLAs. Lead incident response and conduct thorough post-mortems to improve system resilience. What We Need To See 8+ years or equivalent industry experience Bachelor's/Master's Degree in Computer Science, Software Engineering, or equivalent experience. Exceptionally strong background in developing CI/CD processes and deployment pipelines using Jenkins. Extensive experience architecting on AWS Cloud and running services such as API Gateway, Lambda, EKS/ECS, RDS, S3, and SQS. Expert-level knowledge of Terraform (including state management, workspaces, and complex module development). Advanced experience with Kubernetes (EKS) and Docker, including orchestration, service meshes, and Helm. Strong proficiency in a scripting language, such as Python, for automation and custom tooling. Strong communication skills. Ways To Stand Out From The Crowd Deep understanding of DNS and CDNs (e.g., Akamai, CloudFront). Demonstrated use of AI tools to improve productivity and the quality of releases. Applies secure-by-design principles across infrastructure, deployment automation, and operational processes. AWS certifications are preferred. , , JR2021184
More at NVIDIA
Related open roles
Senior Site Reliability Engineer, Senior Site Reliability Engineer
Mumbai
HPC Infra Engineer (Hyderabad)
Hyderabad
Compute Cluster SRE Engineer, GPU - HPC (Bengaluru)
Bangalore
Senior Platform and EngOps Engineer - Cluster Operations
Bangalore
Senior DevOps Engineer - E-commerce
Mumbai
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
India