Source description
About the role
As for expenses on this one for travel, they will bill WWT for expenses directly.
Machine Learning Performance Engineer - CUDA Python - - Travel - 50% - 18966
Job Order #: 10540Title: Machine Learning Performance Engineer - CUDA Python
Position/Sheet Notes: MUST HAVE SOME SORT OF RECORDED VIDEO ON WITH NON-USC's FOR INTERNAL USE. This job is open to W2, C2C, USC, GC and other legal status (H1B).
Duration: 6 month contract with the likelihood to extend
Location: Remote but candidates must be willing to travel to different customer sites.
Position Category: Infrastructure
Job Description: *Must be willing to travel 50% of the time
*Must have strong pre-sales abilities i.e. presentation skills, communication skills, etc.
*Must be willing to help train employees and customers
Your part here is optimizing the performance of our models – both training and inference. We care about efficient large-scale training, low-latency inference in real-time systems, and high-throughput inference in research.
Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking, and host- and GPU-level considerations. Zooming in, we also want to ensure our platform makes sense even at the lowest level – is all that throughput actually goodput? Does loading that vector from the L2 cache really take that long?
• An understanding of modern ML techniques and toolsets
• The experience and systems knowledge required to debug a training run's performance end to end
• Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores, and the memory hierarchy
• Debugging and optimization experience using tools like CUDA GDB, NSight Systems, NSight Compute
• Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN, and cuBLAS
• Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization, and asynchronous memory loads
• Background in Infiniband, RoCE, GPUDirect, PXN, rail optimization, and NVLink, and how to use these networking technologies to link up GPU clusters
• An understanding of the collective algorithms supporting distributed GPU training in NCCL or MPI
• An inventive approach and the willingness to ask hard questions about whether we're taking the right approaches and using the right tools
More at 3B Staffing