Padmi

Site Reliability Engineering Senior Lead

MumbaiPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Fidelity National Information Services

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Lead Site Reliability Engineer at FIS, you will be responsible for defining, building, and operating always-on, low-latency, and highly secure payment platforms that power large-scale financial transactions. This is a senior technical role where you will operate at the intersection of distributed systems engineering, cloud platforms, and reliability architecture. Your key responsibilities will include: - Owning and driving reliability outcomes at scale for real-time, distributed payment and transaction processing platforms with strict SLAs, SLOs, and regulatory requirements. - Defining reliability architecture and standards across services, platforms, and infrastructure to shape how systems are designed, deployed, observed, and operated. - Designing and evolving enterprise-grade observability platforms that provide actionable insights into system health, customer experience, and business impact. - Leading and coordinating responses to high-severity production incidents, driving deep root-cause analysis and long-term systemic fixes. - Setting strategy and driving adoption of SRE best practices including error budgets, capacity modeling, resilience testing, and operational readiness. - Architecting automation and self-service platforms to eliminate toil, reduce operational risk, and enable safe, frequent production releases across teams. - Partnering with senior engineering, product, and platform leaders to influence architectural decisions, cloud migration strategy, disaster recovery posture, and long-term platform evolution. - Mentoring senior engineers and technical leads to raise the overall reliability and operational maturity of the organization. To be successful in this role, you should bring: - Deep software engineering expertise in building and operating large-scale, distributed, API-driven systems in production. - Expertise in observability, alerting, and reliability engineering using tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or equivalent ecosystems. - Strong command of cloud platforms and open systems (AWS, Azure, or GCP) including infrastructure-as-code, platform automation, and cloud-native design patterns. - Significant experience running mission-critical systems in regulated environments such as Payments, FinTech, or Banking. - Hands-on experience across Linux, Windows, databases (e.g., Oracle RDBMS), and complex enterprise stacks with strong troubleshooting skills. - Demonstrated leadership in incident management, continuous reliability improvement, and influencing behavior and standards across teams. - Ability to operate effectively at Staff level scope, solving ambiguous problems, making trade-offs, and driving alignment across multiple teams and stakeholders. Additionally, the following would be an added advantage: - Strong automation and scripting skills using Python, Bash, Ansible, or similar tools. - Experience in building or scaling CI/CD platforms and release automation in high-risk production environments. - Prior ownership of reliability strategy or platform initiatives spanning multiple teams or business units. As a Senior Lead Site Reliability Engineer at FIS, you will be responsible for defining, building, and operating always-on, low-latency, and highly secure payment platforms that power large-scale financial transactions. This is a senior technical role where you will operate at the intersection of distributed systems engineering, cloud platforms, and reliability architecture. Your key responsibilities will include: - Owning and driving reliability outcomes at scale for real-time, distributed payment and transaction processing platforms with strict SLAs, SLOs, and regulatory requirements. - Defining reliability architecture and standards across services, platforms, and infrastructure to shape how systems are designed, deployed, observed, and operated. - Designing and evolving enterprise-grade observability platforms that provide actionable insights into system health, customer experience, and business impact. - Leading and coordinating responses to high-severity production incidents, driving deep root-cause analysis and long-term systemic fixes. - Setting strategy and driving adoption of SRE best practices including error budgets, capacity modeling, resilience testing, and operational readiness. - Architecting automation and self-service platforms to eliminate toil, reduce operational risk, and enable safe, frequent production releases across teams. - Partnering with senior engineering, product, and platform leaders to influence architectural decisions, cloud migration strategy, disaster recovery posture, and long-term platform evolution. - Mentoring senior engineers and technical leads to raise the overall reliability and operational maturity of the organization. To be successful in this role, you should bring: - Deep software engineering expertise in building and operating large-scale, distributed, API-driven systems in production. - Expert

One address, no account. We’ll tell you when matching roles go live.

More at Fidelity National Information Services

Related open roles

View all roles