Source description
About the role
About the Team
The DCS team supports the company's fast growth by building and operating hyperscale datacenters. The team manages the end to end lifecycle of server fleet, providing cloud solutions and various infrastructure services ensuring that they are scalable and are reliable.
- Design, build, scale, and operate ByteDance’s global infrastructure, including large-scale systems spanning public and private clouds.
- Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
- Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards.
- Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
- Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.
More at ByteDance
Related open roles
Site Reliability Engineer - Video Infrastructure
San Francisco Bay Area
Site Reliability Engineer (Cloud) - Infrastructure Engineering
Singapore
Datacenter Operations Engineer
Saudi Arabia
Site Reliability Engineer - Traffic Infrastructure
Singapore
Traffic Access Architectural SRE Graduate (Traffic Infrastructure) - 2026 Start (BS/MS)
Singapore
Site Reliability Engineer, Traffic Solution - System Service Global
Singapore