Padmi

Senior Research Scientist - Machine Learning System

San Francisco Bay AreaPosted 1 month ago
Software engineeringUnspecified
Apply at ByteDance

Opens the source posting on joinbytedance.com

Source description

About the role

View original

The Machine Learning (ML) System sub-team combines system engineering and the art of machine learning to develop and maintain massively distributed ML training and Inference system/services around the world, providing high-performance, highly reliable, scalable systems for LLM/AIGC/AGI

In our team, you'll have the opportunity to build the large scale heterogeneous system integrating with GPU/NPU/RDMA/Storage and keep it running stably and reliably, enrich your expertise in coding, performance analysis and distributed system, and be involved in the decision-making process. You'll also be part of a global team with members from the United States, China and Singapore working collaboratively towards unified project direction.

Responsibilities

  • Responsible for developing and optimizing LLM inference framework. - Responsible for GPU and CUDA Performance optimization to create an industry-leading high-performance LLM inference engine.

One address, no account. We’ll tell you when matching roles go live.

More at ByteDance

Related open roles

View all roles