Padmi

Site Reliability Engineer - ARK Large Model Platform (Singapore)

SingaporePosted 1 month ago
Infrastructure And DatabasesUnspecified
Apply at ByteDance

Opens the source posting on joinbytedance.com

Source description

About the role

View original

About the Team

The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.

In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.

Responsibilities

Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world.

  • Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices.
  • Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance.
  • Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.

One address, no account. We’ll tell you when matching roles go live.

More at ByteDance

Related open roles

View all roles