Source description
About the role
Description Technology/Service is responsible for delivering the business vision and strategy at a global level, focusing on achieving consistent operational excellence, platform reliability, automation, and client/user satisfaction through industrialization, technology modernization, and engineering excellence.The role requires a highly experienced Senior Site Reliability Engineer (SRE) / DevOps Engineering Lead with strong expertise in Cloud, Automation, Platform Engineering, Reliability Engineering, Critical Incident Management, and Data Platform Operations.The ideal candidate should possess strong technical leadership capabilities, an automation-first mindset, and the ability to drive complex incidents toward resolution while ensuring platform stability and operational excellence.What well offer you As part of our flexible scheme, here are just some of the benefits that youll enjoy Best in class leave policyGender neutral parental leaves100% reimbursement under childcare assistance benefit (gender neutral)Sponsorship for Industry relevant certifications and educationEmployee Assistance Program for you and your family membersComprehensive Hospitalization Insurance for you and your dependentsAccident and Term Life InsuranceComplementary Health screening for 35 yrs. and aboveYour key responsibilities Digital & Technology Strategy Create and drive technology transformation initiatives aligned with organizational goals.Identify opportunities to modernize platforms through cloud adoption, automation, and engineering best practices.Act as a technology change agent promoting DevOps, SRE, and automation culture.Apply industry-leading practices and emerging technology trends to improve platform reliability and operational efficiency.CI/CD & Release Engineering Design, implement, and optimize CI/CD pipelines using Jenkins, GitHub Actions, and Azure DevOps.Standardize build, testing, and deployment frameworks across applications and platforms.Integrate automated testing practices including unit, integration, performance, security, and regression testing.Drive trunk-based development and modern DevOps engineering practices.Improve deployment reliability through release orchestration and deployment automation.Cloud & Infrastructure Automation Build and manage infrastructure using Infrastructure as Code (Terraform preferred).Support and optimize cloud-native deployments across Azure, AWS, and GCP environments.Implement containerization strategies utilizing Docker and Kubernetes platforms (AKS, EKS, GKE).Define automated environment provisioning, capacity planning, scaling, and cost optimization strategies.Implement secure, scalable, and resilient cloud infrastructure architectures.Platform Stability & Reliability Engineering Improve platform resilience, availability, observability, and production stability.Establish and drive SRE practices across the platform landscape.Strong understanding of SRE principles including Customer/User Journeys (CUJ), SLI, SLO, SLA, Error Budgets, and Non-Functional Requirements (NFRs).Implement enterprise monitoring and observability solutions using Splunk, New Relic, Prometheus, Grafana, Dynatrace, Elastic Stack, App Insights, or equivalent platforms.Lead Root Cause Analysis (RCA) activities and define preventive and corrective action plans.Design proactive monitoring, alerting, and reliability controls that reduce operational risks and improve customer experience.Data Platform Reliability & Engineering Support and manage large-scale Data Platforms and business-critical data ecosystems.Strong understanding of end-to-end data workflows, data lineage, and data lifecycle management.Ability to analyze and troubleshoot upstream and downstream data dependencies across multiple platforms.Experience supporting batch and real-time data processing environments.Understanding of ETL processes, data ingestion, transformation, reconciliation, and data quality validation.Monitor and manage data pipeline health, data delivery reliability, and operational performance.Proactively identify bottlenecks, data quality issues, and platform risks impacting business operations.Collaborate with Data Engineering and Application teams to enhance data reliability and operational excellence.ITIL, Incident & Service Management Strong understanding of ITIL practices including:Incident ManagementMajor Incident ManagementProblem ManagementChange ManagementService Request ManagementService Level ManagementLead critical production incidents and coordinate technical war-room activities.Drive incidents toward resolution through effective technical leadership, stakeholder management, and cross-team collaboration.Ensure timely service restoration while minimizing business impact.Conduct Post Incident Reviews (PIR) and Problem Management activities.Identify recurring issues and implement permanent corrective actions.Drive continuous service improvement initiatives based on operational learnings.Ensure adherence D
More at E901 DWS India Private Limited, Bangalore Branch