Source description
About the role
This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs. Responsibilities include:
Develop analytical performance models for GPU kernels and inference workloads.
Build and validate a simulator to estimate theoretical hardware performance limits.
Compare measured kernel performance against architectural peak throughput.
Identify performance bottlenecks in compute, memory, communication, and scheduling.
Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
Investigate PTX and SASS code generation to understand low-level execution behavior.
Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
Design profiling methodologies for Hopper and Blackwell architectures.
Document findings and provide actionable recommendations for performance improvements.
More at Institute of Foundation Models
Related open roles
Eval360 - Error Analysis Engineer
San Francisco Bay Area · Onsite
Machine Learning Engineer – World Model
San Francisco Bay Area · Onsite
Research Engineer - Speech/Audio Machine Learning
Paris · Onsite
Machine Learning Infrastructure Engineer
United States · Onsite
Machine Learning Engineer – World Modeling
Dubai · Onsite
Machine Learning Engineer
Dubai · Onsite
