Padmi
NatWest logo
NatWest

retail banking · commercial banking

Site Reliability Engineer, AVP

ChennaiPosted 1 month ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at NatWest

Opens the source posting on shine.com

Source description

About the role

View original

Join us as a Site Reliability Engineer Youll manage the provision of stable, resilient, reliable applications with the end goal of minimising disruption to Customer & Colleague Journeys (CCJ)Well look to you to identify and automate manual tasks and implement observability solutions, ensuring a thorough understanding of CCJ across applicationsThis is a great chance to work in a supportive environment with opportunities to advance your personal and career developmentWe're offering this role at associate vice president level What you'll do As a Site Reliability Engineer, youll collaborate with feature teams to understand application changes, participate in delivery activities, and address production issues to assist in the delivery of change that does not negatively affect the customer experience. You'll contribute to site reliability operations which will include production support, incident response, on-call rota, toil reduction, and application performance. You'll also proactively lead improvement to release quality into production and provide highly available, performing, and secure production systems. Other responsibilities will include: Delivering automation solutions to minimise and eliminate manual tasks associated with maintaining and supporting the applicationsEnsuring in-depth understanding of the full tech stack on which the application resides and depends onIdentifying alerting and monitoring requirements for an application, based on sound understanding of customer journeysEvaluating the resilience of the end-to-end tech stack on which the applications depend, and addressing weaknessesSeeking to reduce frequency of hand-offs in the end-to-end resolution of customer-impacting incidents The skills you'll need To succeed in this role, youll need experience of supporting live production services serving customer journeys with a demonstrable knowledge of ITIL processes and IT Security principles along with tools and techniques to prevent compliance breaches. You'll have hands on experience with Azure Cloud and full-stack observability using tools such as Log Analytics, Application Insights, and Grafana. Youll also need: 8+ years of experience in Production Support, Site Reliability Engineering (SRE), IT Operations, or Application Support within large-scale enterprise environmentsLead incident management, major outage resolution, stakeholder communication, and root cause analysis (RCA) to drive permanent fixesEnsure end-to-end production stability, availability, and performance, meeting agreed SLAs/SLOs and customer service expectationsDrive automation, operational excellence, resilience, capacity planning, change governance, and continuous service improvement initiativesEstablish and maintain monitoring, observability, alerting, dashboards, and service health metrics for proactive issue detection and resolutionStrong verbal and written communication skills .

One address, no account. We’ll tell you when matching roles go live.

More at NatWest

Related open roles

View all roles