Padmi
NatWest logo
NatWest

retail banking · commercial banking

Site Reliability Engineer (AWS & Kubernetes), VP

IndiaPosted 3 months ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at NatWest

Opens the source posting on shine.com

Source description

About the role

View original

As a Site Reliability Engineer, you will play a crucial role in improving, driving, and embedding non-functional and operational characteristics such as availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning of products and services. Your responsibilities will include significant stakeholder interaction, collaborating with engineers to ensure a principled approach to deliver change in a safe and secure manner. This opportunity allows you to join an inclusive team with a collaborative ethos and a commitment to innovation and career development. The role is offered at vice president level. Key Responsibilities: - Designing and operating highly resilient AWS-based Kubernetes platforms aligned to enterprise standards - Owning and improving production reliability, availability, and SLA/SLO frameworks - Leading incident management, escalation, and 24/7 on-call practices, including post-incident reviews - Embedding SRE principles such as error budgets, toil reduction, and reliability engineering into delivery teams - Implementing infrastructure and platform automation using Terraform and GitOps methodologies - Driving self-healing, auto-scaling, and failure recovery mechanisms using tools such as Karpenter - Building secure, scalable networking and service communication (e.g. Cilium) - Defining and operating observability platforms using Grafana, Prometheus, Loki, Tempo - Partnering with DevOps and engineering teams to ensure production readiness and operational excellence - Leading complex troubleshooting across distributed systems and cloud-native environments - Developing reusable golden paths, operational runbooks, and reliability patterns - Ensuring platforms meet regulatory, security, and operational risk requirements - Using data, SLIs, and metrics to drive continuous improvement and proactive reliability enhancements Qualifications Required: - Deep expertise managing production systems on AWS and Kubernetes (EKS) - Strong experience in 24/7 support models, incident management, and on-call leadership - Advanced knowledge of SRE principles (SLIs, SLOs, error budgets, toil reduction) - Proficiency in Terraform, GitOps, and cloud automation practices - Hands-on experience with GitLab CI/CD and Argo CD - Strong understanding of Kubernetes networking, security, and service mesh technologies, ideally Cilium - Experience scaling infrastructure using Karpenter and auto-scaling strategies - Expertise in observability tooling (Grafana, Prometheus, Loki, Tempo) - Proven ability to troubleshoot and resolve complex, cross-system production issues - Experience operating in regulated or high-security environments - Strong leadership, mentoring, and stakeholder engagement capabilities - Ability to balance reliability, risk, and delivery in a fast-paced environment Please note that the job posting closes on 16/06/2026. As a Site Reliability Engineer, you will play a crucial role in improving, driving, and embedding non-functional and operational characteristics such as availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning of products and services. Your responsibilities will include significant stakeholder interaction, collaborating with engineers to ensure a principled approach to deliver change in a safe and secure manner. This opportunity allows you to join an inclusive team with a collaborative ethos and a commitment to innovation and career development. The role is offered at vice president level. Key Responsibilities: - Designing and operating highly resilient AWS-based Kubernetes platforms aligned to enterprise standards - Owning and improving production reliability, availability, and SLA/SLO frameworks - Leading incident management, escalation, and 24/7 on-call practices, including post-incident reviews - Embedding SRE principles such as error budgets, toil reduction, and reliability engineering into delivery teams - Implementing infrastructure and platform automation using Terraform and GitOps methodologies - Driving self-healing, auto-scaling, and failure recovery mechanisms using tools such as Karpenter - Building secure, scalable networking and service communication (e.g. Cilium) - Defining and operating observability platforms using Grafana, Prometheus, Loki, Tempo - Partnering with DevOps and engineering teams to ensure production readiness and operational excellence - Leading complex troubleshooting across distributed systems and cloud-native environments - Developing reusable golden paths, operational runbooks, and reliability patterns - Ensuring platforms meet regulatory, security, and operational risk requirements - Using data, SLIs, and metrics to drive continuous improvement and proactive reliability enhancements Qualifications Required: - Deep expertise managing production systems on AWS and Kubernetes (EKS) - Strong experience in 24/7 support model

One address, no account. We’ll tell you when matching roles go live.

More at NatWest

Related open roles

View all roles