Source description
About the role
We are looking for a Senior SRE Engineer to join our collaborative team. Our team is building the platform fueling digital transformation. We take a cloud-first approach to deliver customer-centric experiences and platform web services that are used by product teams and partners. As a Senior SRE Engineer, you'll be a vital part of the Engagement Team for Digital Transformation. This team enables and aids product teams and partners by adopting and integrating cloud services with a customer-centric ideology always in mind. You will be an expert in cloud services, building, proving out, and communicating the value of the platform that is enabling Digital Transformation. Responsibilities Design and implement high-availability and disaster recovery strategies for critical workloads Build, maintain, and evolve cloud infrastructure using Infrastructure-as-Code practices Develop and manage CI/CD pipelines to enable reliable and efficient delivery Implement and maintain monitoring, observability, and alerting across the platform Lead cross-functional reliability initiatives and platform-wide automation projects Ensure networking, security, and identity/access management best practices in the cloud Collaborate with product teams and partners to adopt and integrate cloud services Optimize infrastructure costs by applying FinOps practices and cost optimization frameworks Troubleshoot production issues, perform root cause analysis, and drive improvements Mentor engineers and influence reliability standards across teams Requirements 5+ years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles Deep AWS expertise (EC2, S3, RDS, IAM, VPC, Lambda, CloudFormation/Terraform, etc.) Strong knowledge of Infrastructure-as-Code (IaC) using Terraform, AWS CDK, or CloudFormation Proven experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or similar Proficiency in containerization and orchestration with Docker, Kubernetes, and ECS or EKS Expertise in monitoring and observability tools (Datadog, New Relic, Prometheus, Grafana, ELK, CloudWatch, etc.) Strong scripting or programming competency in Python, Bash, or Go Sound understanding of networking, security, and identity/access management in the cloud Experience designing high-availability and disaster recovery strategies for critical workloads Excellent communication, problem-solving, and leadership skills with the ability to influence across teams Nice to have AWS or other Cloud Certification (Solutions Architect, DevOps Engineer, etc.) Experience with AIOps, Serverless Architectures, and event-driven systems Familiarity with FinOps practices and cost optimization frameworks Experience with SaaS monitoring tools such as Sumo Logic, PagerDuty, or similar Exposure to Atlassian tools (Jira, Confluence, Bitbucket) Familiarity with SQL/NoSQL databases .
More at EPAM Systems
Related open roles
Java Full-Stack Developer (Angular) - Hyderabad
Hyderabad · Hybrid
Java Full-Stack Developer (Angular) - Hyderabad
Hyderabad · Hybrid
Senior Java Full-Stack Developer (Angular)
Bangalore · Hybrid
Java Full-Stack Developer (Angular)
Bangalore · Hybrid
Data Technology Consultant
Mumbai
Senior Data Software Engineer
Mumbai