Source description
About the role
Computer Scientist II The Developer Platforms team at Adobe is on a mission to build, deploy, operate, and scale the core systems that power Adobe solutions used by millions of customers worldwide. We are looking for experienced Senior Platform / Site Reliability Engineers (SREs) to help design and operate highly reliable, scalable infrastructure that will shape the future of Creative Cloud and Adobes Digital Media business. This role focuses on cloud infrastructure, Kubernetes platforms, service mesh, and developer enablement , with a strong emphasis on automation, resilience, and platform abstraction . You will act as a force multiplier, enabling engineering teams to build and operate services efficiently through robust platform capabilities. What you'll do: Drive adoption of Infrastructure as Code (Terraform) and enforce platform engineering guidelines and standards Define, implement, and continuously improve SLIs, SLOs, and error budgets for Project Graph HTTP APIs and asynchronous compute platforms Standardise and optimise Kubernetes-based deployments , scheduling, and resource management for performance and cost efficiency Build and enable self-service infrastructure platforms that empower product teams and improve developer productivity Lead the adoption of intelligent software solutions (AI agents and MCPs) to reduce operational toil and improve automation at scale Define and enforce platform-wide observability standards , covering metrics, logging, tracing, and alerting Design, build, and maintain robust observability systems to enable rapid detection, diagnosis, and resolution of issues Lead incident response , drive blameless postmortems , and ensure effective follow-through to prevent recurrence Improve the reliability, scalability, and performance of asynchronous job scheduling systems built on Kubernetes and Postgres Design, maintain, and optimise CI/CD pipelines to ensure fast, safe, and reliable software delivery Own data protection and resilience strategies , including backups, restoration testing, and disaster recovery planning Architect and implement cloud infrastructure (AWS) aligned to reliability, scalability, performance, and cost optimisation goals Reduce operational toil through automation, tooling, and platform improvements , partnering closely with developers to build reliability by design Establish and drive guidelines in incident management, reliability engineering, and operational excellence Contribute to evaluating system demand, load testing, and performance engineering to ensure systems scale predictably Ensure security and compliance guidelines are embedded in platform build and operations (e.g., least privilege, secrets management, secure configurations) Participate in an on-call rotation , supporting production systems and driving continuous improvement in operational readiness What you need to succeed: Bachelors degree in Computer Science (or equivalent practical experience) 810 years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure, or Backend Development with a strong operational and production focus AIFirst & Automation Mindset Demonstrated AI-first and automation-first attitude across SRE functions Hands-on experience or exposure to AI Agents, MCPs, and AIOps to drive operational efficiency and reduce toil Proven ability to identify, prioritise, and eliminate operational toil through automation and intelligent tooling Cloud, Containers & Platform Engineering Deep expertise in Kubernetes in production, including scaling, performance tuning, resolving challenges, and workload optimisation Strong experience with containerisation (Docker),Argo and modern deployment patterns Hands-on experience designing and operating cloud-native systems on AWS Proficiency in Infrastructure as Code (Terraform) and cloud automation Programming & Systems Engineering Strong programming skills in Golang and/or Python , with experience building production-grade systems and tooling Experience with Node.js/TypeScript or similar backend technologies is a plus Solid understanding of Linux systems, networking fundamentals, REST APIs, and distributed systems design CI/CD & Developer Productivity Strong experience with CI/CD pipelines and tooling (e.g., Argo, Jenkins, Github Actions, CircleCI) Ability to design and maintain scalable, reliable build and deployment systems Observability & Reliability Engineering Strong expertise in observability practices , including metrics, logging, tracing, and alerting Hands-on experience with tools such as Prometheus, Alertmanager, OpenTelemetry, Jaeger, New Relic , or similar Experience defining and operating SLIs, SLOs, and error budgets Experience with incident response, production operations , and driving blameless postmortems Data & Persistence Layer Expertise Experience operating production-grade databases (Postgres, MySQL/Aurora, Redis, or similar) Strong understanding of data protection, backups, resilience, and .
More at Adobe