Padmi

Senior AWS Site Reliability Engineer

United States · OnsitePosted 2 months ago
Infrastructure And DatabasesUnspecified
Apply at Prophecy Technologies

Opens the source posting on prophecytechs.com

Source description

About the role

View original

Client: TCS | Engagement: Full-time | Work mode: ONSITE | Experience: 10+ | Publisher job id: JOB-000085

Role Overview

We are seeking a highly experienced Senior Site Reliability Engineer (SRE) / Application Reliability Engineer with over 10 years of expertise in incident management, system reliability, and enterprise application support, specifically with AWS knowledge. This role is critical for ensuring high availability, operational stability, and continuous improvement of vital financial and ERP systems within a 24x7 production environment. The ideal candidate will possess strong hands-on experience in monitoring, troubleshooting, root cause analysis, and supporting both cloud-based and on-premise enterprise platforms. Key Responsibilities: Ensure high availability and reliability of enterprise applications in a 24x7 production environment. Monitor applications, batch jobs, and workflows to maintain operational continuity. Lead and manage major incidents (P1/P2) and drive resolution to minimize business impact. Perform root cause analysis (RCA) and implement preventive measures. Ensure adherence to SLA/SLO and ITIL-based incident, problem, and change management processes. Design and maintain monitoring dashboards. Implement proactive alerting and improve system observability. Diagnose and resolve application and data-related issues using SQL queries and log analysis. Provide backend validation and technical support across distributed environments. Support release deployments, change validation, and post-deployment activities. Participate in disaster recovery testing and release readiness validation. Collaborate with infrastructure, DBA, and development teams to resolve technical issues. Create and maintain operational documentation, runbooks, and knowledge base articles. Required Skills: Site Reliability Engineering (SRE) and Application Support Incident & Problem Management Root Cause Analysis (RCA) SLA / SLO Compliance Batch Monitoring & Scheduling ITIL Framework CI/CD Tools: GitHub Cloud Platforms: AWS (EC2, S3, VPC) Databases: Oracle SQL Server Languages: SQL SQR Basic Java Ticketing Tools: ServiceNow Jira Operating Systems: UNIX Linux Windows Qualifications: 10+ years of experience in Application Support / Reliability Engineering roles. Strong experience in BFSI or enterprise application environments. Proven track record in managing production support operations and high-severity incidents.

One address, no account. We’ll tell you when matching roles go live.

More at Prophecy Technologies

Related open roles

View all roles