Source description
About the role
We are looking for an experienced Lead SRE Engineer with strong Linux and monitoring expertise who can operate independently without supervision. In this role, you will join the project, assess the current infrastructure and observability, identify areas for improvement, and drive initiatives to enhance them. Your mission will be to bring visibility, monitoring, and control to a complex file transfer platform. Responsibilities Build end-to-end observability, including transfer tracking, queue monitoring, and latency metrics Improve alerting by reducing noise and increasing signal Instrument legacy and USS systems Design and build dashboards for incidents and system health Partner with the AI team to enable data-driven insights Assess current infrastructure and observability, and identify areas for improvement Drive improvement initiatives independently without direct supervision Requirements 8+ years of experience as an SRE, Platform, or Linux Engineer At least 1 year of relevant leadership experience Strong background in Linux Proficiency in Splunk, Dynatrace, Prometheus, and Grafana Skills in Python and Bash Expertise in debugging distributed systems Skills in performance tuning Capability to work autonomously, assess environments, and drive improvements without supervision Proficiency in English at an Upper-Intermediate level (B2) or higher Nice to have Exposure to Mainframe and USS environments Familiarity with file transfer and batch systems We are looking for an experienced Lead SRE Engineer with strong Linux and monitoring expertise who can operate independently without supervision. In this role, you will join the project, assess the current infrastructure and observability, identify areas for improvement, and drive initiatives to enhance them. Your mission will be to bring visibility, monitoring, and control to a complex file transfer platform. Responsibilities Build end-to-end observability, including transfer tracking, queue monitoring, and latency metrics Improve alerting by reducing noise and increasing signal Instrument legacy and USS systems Design and build dashboards for incidents and system health Partner with the AI team to enable data-driven insights Assess current infrastructure and observability, and identify areas for improvement Drive improvement initiatives independently without direct supervision Requirements 8+ years of experience as an SRE, Platform, or Linux Engineer At least 1 year of relevant leadership experience Strong background in Linux Proficiency in Splunk, Dynatrace, Prometheus, and Grafana Skills in Python and Bash Expertise in debugging distributed systems Skills in performance tuning Capability to work autonomously, assess environments, and drive improvements without supervision Proficiency in English at an Upper-Intermediate level (B2) or higher Nice to have Exposure to Mainframe and USS environments Familiarity with file transfer and batch systems
More at EPAM Systems
Related open roles
Java Full-Stack Developer (Angular) - Hyderabad
Hyderabad · Hybrid
Java Full-Stack Developer (Angular) - Hyderabad
Hyderabad · Hybrid
Senior Java Full-Stack Developer (Angular)
Bangalore · Hybrid
Java Full-Stack Developer (Angular)
Bangalore · Hybrid
Data Technology Consultant
Mumbai
Senior Data Software Engineer
Mumbai