Padmi
Zensar logo
Zensar

AI-led digital transformation · Cloud engineering and DevOps

SRE (AWS or Azure)

MumbaiPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Zensar

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: The Software Engineer / Site Reliability Engineer (SRE) will play a critical role in driving reliability, scalability, and performance for the Banking Solutions, Payments, and Capital Markets platforms. This role blends core SRE principles, performance engineering, and service health management to support large-scale, mission-critical systems. The ideal candidate will help modernize platforms through automation-first practices, data-driven reliability metrics, and proactive performance optimization, ensuring exceptional customer experience and business continuity in a highly regulated environment. Key Responsibilities: - Design, implement, and operate highly available, resilient, and scalable systems aligned with SRE best practices. - Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to balance reliability and delivery velocity. - Build and maintain service health dashboards to provide real-time visibility into platform stability and customer experience. - Reduce toil through extensive automation of operational workflows, alerts, and remediation activities. - Design and maintain end-to-end monitoring and observability solutions covering infrastructure, applications, APIs, and user journeys. - Implement advanced alerting strategies to reduce noise and improve mean time to detect (MTTD) and mean time to resolution (MTTR). - Leverage metrics, logs, and traces to drive root cause analysis and proactive incident prevention. - Enable reliability reporting for stakeholders using SLO compliance and service health metrics. - Lead performance engineering initiatives, including load testing, stress testing, endurance testing, and capacity validation. - Identify performance bottlenecks across application, middleware, database, and infrastructure layers. - Conduct capacity planning and performance tuning to support business growth and peak traffic scenarios. - Partner with development and QA teams to embed performance testing into CI/CD pipelines. - Lead and participate in incident response activities, including triage, mitigation, recovery, and post-incident reviews. - Drive blameless post-mortems and ensure corrective actions are tracked to completion. - Participate in on-call rotations, providing 24x7 support for critical production systems. - Continuously improve operational readiness and resilience. - Design and manage deployment pipelines, configuration management, and environment consistency across lower and production environments. - Implement Infrastructure as Code (IaC) practices for repeatable and secure cloud provisioning. - Collaborate with DevOps teams to improve deployment reliability, rollback mechanisms, and release safety. - Develop and test disaster recovery plans, backup strategies, and failover mechanisms. - Work closely with Development, QA, DevOps, Security, and Product teams to align on reliability and performance goals. - Ensure platforms meet security, compliance, and regulatory requirements common in financial services. - Act as a reliability and performance advocate throughout the SDLC. Qualifications Required: - Strong experience in Core SRE practices, including reliability engineering, incident management, and automation. - Proven hands-on experience in Performance Engineering / Performance Testing for large-scale distributed systems. - Deep understanding and implementation experience with SLI / SLO / Error Budget frameworks. - Proficiency in cloud platforms (AWS, Azure, or Google Cloud). - Hands-on experience with containerization and orchestration (Docker, Kubernetes). - Strong background in monitoring, observability, and logging tools such as Prometheus, Grafana, Datadog, Splunk, ELK Stack. - Experience with CI/CD pipelines (Jenkins, GitLab CI/CD, Azure DevOps). - Proficiency in scripting and automation using Python, Bash, Terraform, Ansible. - Strong troubleshooting skills across application, infrastructure, and network layers. - Experience designing and running incident response and post-mortem reviews. - Ownership mindset with accountability for service reliability and customer outcomes. - Excellent communication, collaboration, and stakeholder management skills. Role Overview: The Software Engineer / Site Reliability Engineer (SRE) will play a critical role in driving reliability, scalability, and performance for the Banking Solutions, Payments, and Capital Markets platforms. This role blends core SRE principles, performance engineering, and service health management to support large-scale, mission-critical systems. The ideal candidate will help modernize platforms through automation-first practices, data-driven reliability metrics, and proactive performance optimization, ensuring exceptional customer experience and business continuity in a highly regulated environment. Key Responsibilities: - Design, implement, and operate highly available, resilient, and scalable systems aligned with SRE best practices. - Define

One address, no account. We’ll tell you when matching roles go live.

More at Zensar

Related open roles

View all roles