Padmi

Site Reliability Engineer 2

Bangalore · HybridPosted 2 months ago
Infrastructure And DatabasesMid-level
Apply at Data Sutram

Opens the source posting on naukri.com

Source description

About the role

View original

About the role In this role, you will be a key contributor to production infrastructure and reliability for high-scale, regulated systems. You will work across cloud infrastructure, Kubernetes platforms, observability, incident management, CI/CD, and internal tooling helping engineering teams ship confidently while meeting stringent availability and scalability requirements. This role expects strong hands-on execution, a growing sense of architectural judgment, and the ability to independently own reliability and platform tasks end-to-end, with guidance from senior engineers on larger initiatives. Key Responsibilities Build and operate AWS-based production infrastructure spanning networking, compute, storage, and observability. Support and operate Kubernetes (EKS) workloads, including deployments, scaling, and day-to-day operational health. Build and maintain infrastructure automation using Terraform and GitOps-driven CI/CD pipelines (ArgoCD, GitHub Actions), improving speed and reliability. Design and build internal tools and self-service platforms from ephemeral environments to testing frameworks — that speed up other engineering teams' workflows. Contribute to fault-tolerant, highly available architectures across multiple AZs and assist in executing DR processes against defined RPO/RTO targets. Build and maintain metrics, alerting, and logging pipelines for the services and platforms you own. Participate in incident response and blameless post-mortems, and drive the resulting fixes and toil-reduction work. Work with VPCs, routing, security groups, and ALB/NLB configurations, and partner with Security on defence-in-depth using Cloudflare, firewalls, IAM, and AWS security primitives. Debug Linux, networking, and system performance issues under real production load. Build AI-native solutions using tools like Cursor and Claude to improve operational efficiency and SDLC speed. Qualifications That We Are Looking For 3–6 years of experience in SRE / DevOps / Platform / Infrastructure Engineering roles. Hands-on experience with AWS (EC2, ASG, ALB/NLB, IAM, VPC, networking). Working experience operating Kubernetes (EKS) in production environments. Good understanding of networking fundamentals (routing, DNS, load balancing, firewalls). Experience with Terraform and infrastructure-as-code practices. Exposure to CI/CD tooling (GitHub Actions, Jenkins, ArgoCD, or Spinnaker) and interest in developer productivity problems. Solid Linux fundamentals and production troubleshooting skills. Working knowledge of monitoring, alerting, and logging systems. Proficiency in Golang, Python, or TypeScript (Golang preferred for tooling work). Exposure to service mesh architectures (Istio or similar) is a plus. Experience in Agile/Scrum environments with the ability to independently own and ship tasks end-to-end. B.E/B.Tech/M.E./M.Tech/M.S. from a reputed university with a good academic record. Curiosity to explore cutting-edge technologies and zeal to take ownership of your work; open-source or personal-project tooling experience is a plus.

One address, no account. We’ll tell you when matching roles go live.

More at Data Sutram

Related open roles

View all roles