Source description
About the role
Reporting to: Full Stack Engineer Lead Department: Service Delivery Department (SDD) Domain: Full Stack Engineering Shared Platform Role Overview As the DevOps & Production Engineering - GCC Lead based in our Chennai strategic technology center, you will be the key technical authority and leader responsible for ensuring the 24/7 availability, reliability, and top-tier performance of our next-generation digital banking web and mobile applications across the Asia Pacific region. Operating within a high-stakes, hybrid environment (Cloud and On-Premises), you will champion operational excellence and safeguard customer-facing banking journeys. In this role, you will lead, scale, and mentor the Chennai team of DevOps and production support engineers, working closely with the counterparts in Singapore. This is a high-visibility, hands-on leadership role-you will drive critical incident resolution, orchestrate deep-dive root cause analyses alongside development teams, and transition operations from reactive troubleshooting to long-term proactive engineering solutions. About the Opportunity Strategic Leadership: Own and scale the 24/7 global operational support strategy for a premier digital banking platform from our Chennai technology hub. Hands-on Triage: Actively lead high-severity incident troubleshooting sessions, guiding teams to deliver immediate technical workarounds and permanent fixes. Enterprise Scale Architecture: Oversee complex multi-tier ecosystems featuring web frontends, mobile applications (iOS/Android), extensive microservices, and multiple relational/NoSQL databases across hybrid cloud infrastructures. Cross-Regional Collaboration: Manage and unify diverse, multi-regional onshore and offshore support engineering teams, coordinating seamless follow-the-sun handoffs between Chennai, Singapore, and other tech hubs. SRE & Automation Transformation: Drive the reduction of operational toil by embedding modern Site Reliability Engineering (SRE) and automation principles into traditional support models. Key Responsibilities Incident Management & Hands-On Troubleshooting Crisis Command: Actively lead technical troubleshooting sessions for critical (Severity 1 and 2) production incidents affecting web, mobile, backend services, and critical data layers. Rapid Resolution & Reporting: Ensure rapid restoration of services while maintaining clear, real-time executive communication and business stakeholder updates during major incidents. Root Cause Elimination: Partner closely with Development, Cloud Infrastructure, and DevOps teams to conduct detailed Post-Mortems and Blameless Root Cause Analyses (RCA), tracking temporary workarounds through to permanent software or architectural bug fixes. Team Leadership & Global Operations Follow-the-Sun Governance: Manage and optimize multi-geographical support squads operating across Chennai and Singapore to guarantee seamless, sustainable 24/7 production coverage. Capability Building: Mentor and elevate the technical capabilities of support and DevOps engineers, establishing clear engineering career paths and runbook proficiencies. Operational Readiness: Define, track, and regularly report on critical SLAs, OLAs, and service metrics including MTTR (Mean Time to Resolution) and MTTD (Mean Time to Detection) to senior IT leadership. Hybrid Platform & Database Reliability Hybrid Infrastructure Support: Manage and troubleshoot applications distributed across both traditional On-Premises enterprise datacenters and Cloud native environments (Azure/AWS). Multi-Database Governance: Oversee operational health, query performance, and failover/replication mechanisms across multiple database backends (SQL Server, Oracle, PostgreSQL, NoSQL). Observability Engineering: Drive the design and enhancement of enterprise monitoring, tracing, and logging solutions (e.g., Dynatrace, Datadog, Splunk, ELK) to systematically detect anomalies before they impact end-users. Process Optimization & Automation Toil Elimination: Identify repetitive manual tasks and champion an automation-first approach, developing scripts to automate standard health-checks, recovery procedures, and deployments. ITIL Excellence: Govern the implementation of high-standard ITIL frameworks covering Incident, Problem, Change, and Release management tailored for rapid-deployment environments. Disaster Recovery Strategy: Plan, lead, and execute complex regular Disaster Recovery (DR) and business continuity drills for critical banking channels. Security, Risk & Reporting to: Full Stack Engineer Lead Department: Service Delivery Department (SDD) Domain: Full Stack Engineering Shared Platform Role Overview As the DevOps & Production Engineering - GCC Lead based in our Chennai strategic technology center, you will be the key technical authority and leader responsible for ensuring the 24/7 availability, reliability, and top-tier performance of our next-generation digital banking web and mobile applications