Padmi
Microsoft logo
Microsoft

cloud computing (Azure) · AI and machine learning (Copilot, CoreAI)

Principal Firmware Engineer

Hyderabad · Delhi NCR · OnsitePosted 4 days ago
HardwareStaff+Full Time
Apply at Microsoft

Opens the source posting on apply.careers.microsoft.com

Source description

About the role

View original

Define and drive the technical vision, architecture, and roadmap for scalable system stress and performance characterization frameworks across current and future AI accelerator generations. Design and development of highly-performant validation and stress workloads spanning GPU Compute engines, memory, networking, PCIe, and DMA subsystems. Architect and develop end-to-end validation strategies from pre-silicon environments through post-silicon bring-up, characterization, and production deployment with a goal to identify hardware, firmware, and system-level reliability issues as early as possible Develop end-to-end post-silicon tests and tools for functional and performance scenarios of the system. Develop performance analysis, telemetry, and observability solutions to measure compute utilization, memory bandwidth, network throughput, power, thermal behavior and end-to-end workload performance. Drive adoption of profiling, performance monitoring, and other platform observability technologies to accelerate debugging, tuning, and characterization. Partner closely with Architecture, AI software, Firmware, Silicon Validation, Manufacturing, Performance Engineering, and Cloud Infrastructure teams to influence platform requirements and readiness. Provide technical leadership through design reviews, architecture guidance, and strategic recommendations to engineering leadership. Mentor engineers across workload development, performance optimization, debugging, automation, and validation, while establishing best practices for software quality, CI/CD, telemetry, and large-scale system validation. Drive execution across multiple programs, balancing deep technical contribution with broad organizational impact. BS. or higher in Computer Science, Computer Engineering, Electrical Engineering, or related. 12+ years of experience developing complex software, firmware, system software, or validation frameworks. 8+ years' experience in post-silicon SoC or system validation or diagnostic/microbenchmark/stress test content development. Experience with one or more of these: Accelerators or GPUs, DMA Engines, PCIe, Memory (DDR, HBM), Networking Proven experience in one or more areas: Silicon validation, Firmware development, Platform diagnostics, Performance engineering, Stress and reliability testing Experience debugging issues across hardware, firmware, drivers, SDKs, and applications. Strong leadership and cross-team collaboration skills Experience in Post-silicon tests or tools development for functional and performance scenarios. Knowledge of or experience with AI models/kernels such as GEMM, GPT, Gemma, MoE, Llama, or similar Experience with PyTorch, CUDA, Triton, or accelerator runtime frameworks. Good understanding of AI accelerators, such as GPUs, NPUs, FPGAs, with an understanding of their architectures including Data pipelines, Data formats, Memory hierarchies including HBM/LPDDR, Tensor core architecture etc. Experience with Cuda or GPU or tensor-based programming is a plus+ Experience with pre-silicon environments such as emulation or simulation platforms. Experience developing profiling, and performance analysis tools. Experience in performance engineering, bottleneck identification and performance optimization. Knowledge of power and thermal profiling, TDP/PnP, and PVT characterization. Experience in build systems such as CMake and familiarity with CI/CD systems. Ability to work closely with diverse customers and collaborators across varied disciplines (silicon architecture, FW, SW dev, validation engineers) to reconcile requirements from understanding their needs to resolving their problems.

More at Microsoft

Related open roles

View all roles