Source description
About the role
About The Team
The mission of our AML team is to push the next-generation AI infrastructure and recommendation platform for the ads ranking, search ranking, live & e-Commerce ranking in our company. We also drive substantial impact on core businesses of the company.
Responsibilities
- Responsible for the overall architecture design and implementation of model inference services, building a high-performance, highly available, and scalable enterprise-level inference system for large-parameter, high-complexity AI models, overcoming various architectural challenges in the implementation of complex model inference, and supporting the efficient launch of models across all business scenarios. - Responsible for the R&D and optimization of the core modules of the inference framework, covering core capabilities such as inference engine scheduling, monitoring and alerting, canary release, etc., continuously iterating on the framework performance, and resolving performance bottlenecks, resource bottlenecks, and stability issues in high-concurrency and large-model inference scenarios. - Keep track of the latest inference technologies in the industry, conduct technology selection and innovation in combination with business scenarios, accumulate distributed high-concurrency service architecture solutions, and promote the upgrade and standardization of the team's technical system.
More at ByteDance
Related open roles
Senior Cloud Acceleration Engineer – DPU & AI Infra
Seattle
MLLM Algorithm Engineer Graduate (Search) - 2026 Start (BS/MS)
Singapore
Software Engineer (Treasury) - Global Payments
Singapore
Backend Engineer - Large Model Knowledge System/Content understanding/RAG (ByteDance Singapore)
Singapore
Student Researcher (AI Foundation Model Infrastructure - Seed) - 2027 Start (BS/MS)
Seattle
Research Intern (RDMA/High Speed Network) - 2026 Start (PhD)
San Francisco Bay Area