Source description
About the role
As a Site Reliability Engineer (SRE) at Techdome in Hyderabad / Indore, your primary responsibility will be to ensure the availability, reliability, scalability, and performance of cloud-based production systems across the payments and platform products. This role emphasizes automation, observability, CI/CD, incident management, and leveraging AI tooling to reduce operational toil. Key Responsibilities: - Ensure high availability, performance, and scalability of production systems. - Build automation for deployment, monitoring, and incident response. - Implement observability through metrics, logging, tracing, and alerting tools like Prometheus, Grafana, ELK, and Datadog. - Define and manage SLIs, SLOs, and error budgets. - Develop and maintain CI/CD pipelines and infrastructure as code using Terraform and Ansible. - Lead incident response, conduct root cause analysis, and participate in post-incident reviews. - Perform capacity planning and optimize cloud costs. - Participate in an on-call rotation for production support. Required Skills & Qualifications: - 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Platform Engineer. - Proficiency in cloud platforms like AWS, GCP, or Azure. - Familiarity with containers and orchestration tools such as Docker and Kubernetes. - Experience with infrastructure as code using Terraform and Ansible. - Strong scripting/programming skills in Python, Go, or Bash. - Solid understanding of Linux, networking, and distributed-systems fundamentals. - Experience with CI/CD pipelines like Jenkins, GitHub Actions, GitLab CI, or similar tools. Preferred Skills: - Experience with AI/LLM-powered tooling for ops automation, incident summaries, and alert triage. - Background in payments/fintech production environments. - Knowledge of SLO-driven reliability and on-call process improvement. Techdome is a technology-driven company with over 5 years of experience building products across various industries, including payments and fintech. Joining our team will offer you genuine ownership, rapid growth opportunities, and a collaborative environment where your ideas are valued. The hiring process at Techdome is fast and transparent, utilizing JIA, our in-house AI hiring platform, for consistent application reviews and responses within a working day. The process includes a technical round, a team discussion, and a job offer. As a Site Reliability Engineer (SRE) at Techdome in Hyderabad / Indore, your primary responsibility will be to ensure the availability, reliability, scalability, and performance of cloud-based production systems across the payments and platform products. This role emphasizes automation, observability, CI/CD, incident management, and leveraging AI tooling to reduce operational toil. Key Responsibilities: - Ensure high availability, performance, and scalability of production systems. - Build automation for deployment, monitoring, and incident response. - Implement observability through metrics, logging, tracing, and alerting tools like Prometheus, Grafana, ELK, and Datadog. - Define and manage SLIs, SLOs, and error budgets. - Develop and maintain CI/CD pipelines and infrastructure as code using Terraform and Ansible. - Lead incident response, conduct root cause analysis, and participate in post-incident reviews. - Perform capacity planning and optimize cloud costs. - Participate in an on-call rotation for production support. Required Skills & Qualifications: - 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Platform Engineer. - Proficiency in cloud platforms like AWS, GCP, or Azure. - Familiarity with containers and orchestration tools such as Docker and Kubernetes. - Experience with infrastructure as code using Terraform and Ansible. - Strong scripting/programming skills in Python, Go, or Bash. - Solid understanding of Linux, networking, and distributed-systems fundamentals. - Experience with CI/CD pipelines like Jenkins, GitHub Actions, GitLab CI, or similar tools. Preferred Skills: - Experience with AI/LLM-powered tooling for ops automation, incident summaries, and alert triage. - Background in payments/fintech production environments. - Knowledge of SLO-driven reliability and on-call process improvement. Techdome is a technology-driven company with over 5 years of experience building products across various industries, including payments and fintech. Joining our team will offer you genuine ownership, rapid growth opportunities, and a collaborative environment where your ideas are valued. The hiring process at Techdome is fast and transparent, utilizing JIA, our in-house AI hiring platform, for consistent application reviews and responses within a working day. The process includes a technical round, a team discussion, and a job offer.
More at Techdome
Related open roles
Principal Software Architect - MERN, Healthcare & Payments
Hyderabad
Principal Software Architect - MERN, Healthcare & Payments
Hyderabad
Full Stack Developer (MERN Stack)
Hyderabad
Principal Software Architect - MERN, Healthcare & Payments
Hyderabad
Principal / Staff Architect MERN, Healthcare & Payments
Hyderabad
Principal / Staff Architect MERN, Healthcare & Payments
Hyderabad