Source description
About the role
Identify Training Bottlenecks: Profile and analyze Bird's Eye View (BEV) model training pipelines to pinpoint computational and memory bottlenecks.
Develop Custom Kernels: Design and implement high-performance custom compute kernels using CUDA, Triton, or C++ to accelerate the model training process.
Leverage LLMs for Optimization: Explore and integrate Large Language Models (LLMs) to assist in generating high-performance code and optimizing kernel logic.
Automate Profiling Workflows: Build systems to automate performance profiling and analysis using tools like NVIDIA Nsight and the PyTorch Profiler.
Iterative Performance Tuning: Continuously analyze profiling data generated by both human and LLM-assisted workflows to maximize GPU utilization and reduce training times.
More at PlusAI
Related open roles
Engineering Manager - Simulation and Verification
United States · Onsite
Engineering Technician
San Francisco Bay Area · Onsite
Data Engineer
United States · Hybrid
Software Engineer, Simulation
San Francisco Bay Area · Hybrid
Systems Engineering Intern
San Francisco Bay Area · Onsite
Systems Engineering Intern
San Francisco Bay Area · Onsite
