Source description
About the role
Design, build, and operate highly available, fault-tolerant cloud infrastructure across AWS, GCP, and/or Azure
Architect and maintain scalable CI/CD pipelines and deployment frameworks for enterprise-grade software delivery
Lead infrastructure-as-code adoption and maturity using tools such as Terraform, CloudFormation, and Ansible
Own Kubernetes reliability across multi-cluster environments, including upgrades, scaling, and workload lifecycle management
Establish and evolve observability platforms (metrics, logs, traces) and define SLO/SLI frameworks across teams
Lead incident response for critical outages, drive root cause analysis, and implement preventative improvements
Optimize infrastructure for cost, performance, and scalability, partnering closely with engineering and finance stakeholders
Define and enforce DevOps, reliability, and security best practices across the organization
Partner cross-functionally with engineering, data, QA, security, and IT teams to design resilient systems
Mentor engineers and contribute to technical leadership through design reviews, standards, and knowledge sharing
These responsibilities summarize the role’s primary responsibilities and are not an exhaustive list. They may change at the company’s discretion.
What Success Looks Like in Your First Year
Conduct a comprehensive assessment of the current infrastructure, drive infrastructure-as-code adoption to 95%+ across critical systems, and establish clear health and reliability baselines for the Kubernetes platform
Standardize observability using modern tooling and implement an SLO/SLI framework adopted across multiple product teams, including defined SLAs for critical data systems
Strengthen security and compliance posture across cloud environments by implementing consistent baselines, launching a compliance-as-code framework, and reducing mean time to resolution (MTTR) for production incidents
Define, document, and drive adoption of engineering standards, best practices, and operational guidelines across platform and product teams
Develop and align stakeholders on a forward-looking platform reliability and infrastructure roadmap
Demonstrate measurable mentorship and technical leadership impact across the engineering organization
Evaluate and provide recommendations on emerging infrastructure needs, including support for AI/ML and advanced data workloads
More at GRAIL
Related open roles
Senior Data Scientist #4887
San Francisco Bay Area · Hybrid
Forward Deployed Engineer - Software Development Engineer – Artificial Intelligence (AI) - #4883
San Francisco Bay Area · Hybrid
Equipment Engineer 2 #4778, #4779 (Night Shift 10:00pm - 8:00am, Wed - Sat)
Raleigh–Durham · Onsite
Senior Data Scientist # 4630
San Francisco Bay Area · Hybrid
