Padmi
Adobe logo
Adobe

Creative Cloud · Lightroom Desktop

Senior Infrastructure Developer - AI Platform Engineering

Delhi NCRPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at Adobe

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: You are a Senior Infrastructure Developer with over 10 years of experience, responsible for owning, evolving, and scaling the platform supporting demanding ML training workloads. Your role involves architecting systems, writing production-grade code, leading multi-quarter projects across geo-distributed teams, and setting the reliability standards for an infrastructure crucial for thousands of GPU hours daily. Key Responsibilities: - Architect Kubernetes-native infrastructure for running distributed GPU training jobs at massive scale, focusing on reliability and efficiency - Lead complex, multi-team projects end-to-end, aligning stakeholders across time zones and driving delivery in fast-moving environments - Ship cloud infrastructure through Infrastructure as Code (IaC) tools like Terraform or Pulumi, with the same rigor as application code - Design and maintain deep observability stacks including metrics, distributed tracing, log aggregation, and SLO/SLI frameworks - Develop automation, internal tooling, operators, and platform services in languages like Go, Python, or Rust - Lead incident response, post-mortems, and reliability reviews, driving systemic fixes and setting the on-call culture - Debug and resolve complex cluster networking issues, mentor team members, and raise the technical bar through code reviews and knowledge sharing Qualifications Required: - 10+ years in SRE, platform engineering, or infrastructure roles - Expert-level knowledge of Kubernetes internals, GPU/accelerator training workloads, cloud infrastructure, and Infrastructure as Code - Proficiency in observability tools like Prometheus, Grafana, AlertManager, distributed tracing, log aggregation, and networking fundamentals - Strong coding skills in Go, Python, or Rust, with experience in production-quality code and distributed systems design - Leadership experience in leading cross-functional projects, collaborating across time zones, and effective communication skills Note: Additional details about the company were not provided in the job description. Role Overview: You are a Senior Infrastructure Developer with over 10 years of experience, responsible for owning, evolving, and scaling the platform supporting demanding ML training workloads. Your role involves architecting systems, writing production-grade code, leading multi-quarter projects across geo-distributed teams, and setting the reliability standards for an infrastructure crucial for thousands of GPU hours daily. Key Responsibilities: - Architect Kubernetes-native infrastructure for running distributed GPU training jobs at massive scale, focusing on reliability and efficiency - Lead complex, multi-team projects end-to-end, aligning stakeholders across time zones and driving delivery in fast-moving environments - Ship cloud infrastructure through Infrastructure as Code (IaC) tools like Terraform or Pulumi, with the same rigor as application code - Design and maintain deep observability stacks including metrics, distributed tracing, log aggregation, and SLO/SLI frameworks - Develop automation, internal tooling, operators, and platform services in languages like Go, Python, or Rust - Lead incident response, post-mortems, and reliability reviews, driving systemic fixes and setting the on-call culture - Debug and resolve complex cluster networking issues, mentor team members, and raise the technical bar through code reviews and knowledge sharing Qualifications Required: - 10+ years in SRE, platform engineering, or infrastructure roles - Expert-level knowledge of Kubernetes internals, GPU/accelerator training workloads, cloud infrastructure, and Infrastructure as Code - Proficiency in observability tools like Prometheus, Grafana, AlertManager, distributed tracing, log aggregation, and networking fundamentals - Strong coding skills in Go, Python, or Rust, with experience in production-quality code and distributed systems design - Leadership experience in leading cross-functional projects, collaborating across time zones, and effective communication skills Note: Additional details about the company were not provided in the job description.

One address, no account. We’ll tell you when matching roles go live.

More at Adobe

Related open roles

View all roles