Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Principal Hardware Engineer, AI Systems

San Francisco Bay Area · Seattle · San Diego · OnsitePosted 4 days ago
HardwareStaff+Full TimeH-1B track record
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Serve as the System Technical Lead (STL) and end-to-end technical owner for next-generation AI and GPU platforms from architecture handoff through production deployment and fleet readiness. Partner closely with System Architects to translate product requirements, workload needs, architectural intent, and new technologies into executable system designs, engineering requirements, and development plans. Drive program-level technical execution across hardware, firmware, software, validation, manufacturing, and datacenter infrastructure teams, ensuring alignment to system requirements, architecture specifications, schedule, and quality objectives. Lead cross-functional technical decision making and resolve complex system-level tradeoffs spanning performance, power, thermal, mechanical, reliability, manufacturability, serviceability, cost, and total cost of ownership (TCO). Own technical readiness for key program milestones, design reviews, phase exits, and production releases, ensuring engineering deliverables are complete, integrated, and meet quality expectations. Maintain end-to-end system integrity across electrical, mechanical, thermal, firmware, networking, rack, and datacenter domains, ensuring seamless integration from component to rack and cluster level. Drive alignment across engineering disciplines including Electrical, Mechanical, Thermal, Power, Firmware, System Engineering, Validation, Manufacturing, and Supply Chain teams to deliver a cohesive system solution. Partner with TPMs to establish and manage program technical baselines, assess technical impacts of design changes, identify risks, and drive issue resolution throughout the development lifecycle. Evaluate, de-risk, and enable adoption of new and disruptive technologies, including AI accelerators, advanced memory architectures, liquid cooling solutions, optical interconnects, rack-scale infrastructure, and emerging datacenter technologies. Collaborate with ODMs, silicon suppliers, and ecosystem partners to influence technical direction, resolve critical issues, and ensure successful integration and production readiness. Communicate technical status, risks, mitigation plans, and key decisions to engineering leadership and executive stakeholders while serving as the primary point of accountability for program technical success. Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 7+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience. Proven track record leading cross-functional technical execution across hardware, firmware, software, and datacenter infrastructure. Deep system expertise in power delivery, thermal and liquid cooling, signal integrity, mechanical design, and reliability. Experience bringing high-volume silicon platforms (GPU, SoC, accelerator) from architecture through production ramp. Hands-on experience with PCIe, DDR, Ethernet, BIOS/BMC, and Linux and Windows integration. Experience with datacenter-scale AI systems, including system debug and root cause analysis. Proven ability to evaluate AI systems using performance-per-watt and performance-per-dollar metrics. Clear, concise communicator with the ability to influence technical direction across teams and at senior levels. BS / MS in Electrical/Computer Engineering or equivalent industry experience 10+ years of relevant experience in system (compute, storage, networking, and/or accelerator) level design and/or implementation across the hardware development lifecycle. 10+ years of hands-on experience in server hardware architecture, design, and development with solid understanding of hardware, firmware, and Operating System (OS). Proven experience delivering AI and GPU-based systems to production.

More at Microsoft

Related open roles

View all roles