Source description
About the role
Role Overview: You are a Senior Infrastructure Developer with over 10 years of experience, responsible for owning, evolving, and scaling the platform supporting demanding ML training workloads. Your role involves architecting systems, writing production-grade code, leading multi-quarter projects across geo-distributed teams, and setting the reliability standards for an infrastructure crucial for thousands of GPU hours daily. Key Responsibilities: - Architect Kubernetes-native infrastructure for running distributed GPU training jobs at massive scale, focusing on reliability and efficiency - Lead complex, multi-team projects end-to-end, aligning stakeholders across time zones and driving delivery in fast-moving environments - Ship cloud infrastructure through Infrastructure as Code (IaC) tools like Terraform or Pulumi, with the same rigor as application code - Design and maintain deep observability stacks including metrics, distributed tracing, log aggregation, and SLO/SLI frameworks - Develop automation, internal tooling, operators, and platform services in languages like Go, Python, or Rust - Lead incident response, post-mortems, and reliability reviews, driving systemic fixes and setting the on-call culture - Debug and resolve complex cluster networking issues, mentor team members, and raise the technical bar through code reviews and knowledge sharing Qualifications Required: - 10+ years in SRE, platform engineering, or infrastructure roles - Expert-level knowledge of Kubernetes internals, GPU/accelerator training workloads, cloud infrastructure, and Infrastructure as Code - Proficiency in observability tools like Prometheus, Grafana, AlertManager, distributed tracing, log aggregation, and networking fundamentals - Strong coding skills in Go, Python, or Rust, with experience in production-quality code and distributed systems design - Leadership experience in leading cross-functional projects, collaborating across time zones, and effective communication skills Note: Additional details about the company were not provided in the job description. Role Overview: You are a Senior Infrastructure Developer with over 10 years of experience, responsible for owning, evolving, and scaling the platform supporting demanding ML training workloads. Your role involves architecting systems, writing production-grade code, leading multi-quarter projects across geo-distributed teams, and setting the reliability standards for an infrastructure crucial for thousands of GPU hours daily. Key Responsibilities: - Architect Kubernetes-native infrastructure for running distributed GPU training jobs at massive scale, focusing on reliability and efficiency - Lead complex, multi-team projects end-to-end, aligning stakeholders across time zones and driving delivery in fast-moving environments - Ship cloud infrastructure through Infrastructure as Code (IaC) tools like Terraform or Pulumi, with the same rigor as application code - Design and maintain deep observability stacks including metrics, distributed tracing, log aggregation, and SLO/SLI frameworks - Develop automation, internal tooling, operators, and platform services in languages like Go, Python, or Rust - Lead incident response, post-mortems, and reliability reviews, driving systemic fixes and setting the on-call culture - Debug and resolve complex cluster networking issues, mentor team members, and raise the technical bar through code reviews and knowledge sharing Qualifications Required: - 10+ years in SRE, platform engineering, or infrastructure roles - Expert-level knowledge of Kubernetes internals, GPU/accelerator training workloads, cloud infrastructure, and Infrastructure as Code - Proficiency in observability tools like Prometheus, Grafana, AlertManager, distributed tracing, log aggregation, and networking fundamentals - Strong coding skills in Go, Python, or Rust, with experience in production-quality code and distributed systems design - Leadership experience in leading cross-functional projects, collaborating across time zones, and effective communication skills Note: Additional details about the company were not provided in the job description.
More at Adobe