Padmi

Senior Staff System Engineer, GPU Fleet

BangalorePosted 1 month ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at Coupand

Opens the source posting on shine.com

Source description

About the role

View original

Company Introduction We exist to wow our customers. We know were doing the right thing when we hear our customers say, How did I ever live without Coupang Born out of an obsession to make shopping, eating, and living easier than ever, were collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of whats possible to solve problems and break traditional trade-offs.Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview We are seeking a Sr Staff System Engineer, GPU Fleet for our Coupang Intelligent Cloud (CIC) team, to serve as the senior technical owner for our hyperscale GPU compute infrastructure. In this role, you will define fleet architecture, drive reliability and automation at scale, and lead the operation and evolution of GPU systems supporting largescale AI training and inference workloads. This is a handson, stafflevel individual contributor role with broad technical ownership, high operational impact, and significant crossfunctional influence across hardware, infrastructure, and datacenter operations. CIC builds the infrastructure for abundant intelligence. We partner with leading AI labs, governments, and enterprises to deliver hyperscale GPU compute with high reliability, performance, and efficiency. Our infrastructure supports some of the most demanding AI training and inference workloads in production today. We operate with urgency, deep ownership, and a strong bias toward execution. Reliability, operational excellence, and rigorous systems engineering are core to our business. What You Will Do As a Sr Staff System Engineer, GPU Fleet, you will be the senior technical owner for CICs largescale GPU compute infrastructure. This is a handson senior individual contributor role with fleetlevel responsibility and broad crossfunctional influence. You will define the technical direction for how GPU fleets are architected, operated, automated, and evolved across multiple generations of hardware. Your work will directly affect fleet reliability, operating efficiency, scalability, and customer success. This role does not involve people management, but it carries principallevel scope, autonomy, and decisionmaking authority across infrastructure, hardware, and operations. Key Responsibilities: Fleet Architecture & Technical Ownership Own the endtoend technical architecture of hyperscale GPU fleets, including hardware platform selection, firmware strategy, OS configuration, drivers, networking, and observability. Define and enforce technical standards and best practices for fleet reliability, availability, performance, and operability. Lead major fleetwide initiatives such as new GPU platform bringups, multigeneration hardware transitions, and architectural redesigns. Evaluate tradeoffs across cost, performance, reliability, and timetodeploy, and make technically sound decisions under ambiguity. Reliability, Availability & Performance Set and drive fleetlevel reliability, availability, and performance objectives. Lead rootcause analysis and resolution of complex, systemic failures affecting large portions of the fleet or multiple datacenters. Identify recurring failure patterns and drive longterm fixes spanning hardware, software, automation, and operational processes. Work directly with hardware vendors and partners to resolve platformlevel issues and influence future hardware designs. Automation & Systems Engineering Design and build largescale automation systems for: GPU fleet provisioning and lifecycle management GPU health validation, diagnostics, and certification Automated remediation, recovery, and replacement workflows Eliminate manual operational .

One address, no account. We’ll tell you when matching roles go live.