Source description
About the role
• Monitor health, performance, and availability of large-scale GPU clusters. • Respond to incidents and perform first-level triage. • Support researchers and troubleshoot job failures. • Execute operational runbooks and recovery procedures. • Validate cluster deployments, upgrades, and maintenance activities. • Track infrastructure utilization and operational metrics. • Develop automation and monitoring tools. • Contribute to documentation and reporting.
More at Institute of Foundation Models
Related open roles
Inference Optimization Intern – Performance Modeling
United States · Onsite
Eval360 - Error Analysis Engineer
San Francisco Bay Area · Onsite
AI Research Internship - WM
San Francisco Bay Area · Onsite
Research Scientist, Agentic Data & Benchmarking
San Francisco Bay Area · Onsite
IT Operations Lead
San Francisco Bay Area · Onsite
Senior MLOps Engineer
United Arab Emirates · Onsite
