Padmi

Site Reliability Engineer

HyderabadPosted 2 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at Techdome

Opens the source posting on shine.com

Source description

About the role

View original

As a Site Reliability Engineer (SRE) at Techdome in Hyderabad / Indore, your role will be crucial in ensuring the availability, reliability, scalability, and performance of cloud-based production systems for our payments and platform products. You will primarily focus on automation, observability, CI/CD, incident management, and leveraging AI tooling to reduce operational toil. Key Responsibilities: - Ensure high availability, performance, and scalability of production systems. - Build automation for deployment, monitoring, and incident response. - Implement observability through metrics, logging, tracing, and alerting tools like Prometheus, Grafana, ELK, and Datadog. - Define and manage SLIs, SLOs, and error budgets. - Develop and maintain CI/CD pipelines and infrastructure as code using Terraform and Ansible. - Lead incident response, conduct root cause analysis, and participate in post-incident reviews. - Conduct capacity planning and optimize cloud costs. - Engage in an on-call rotation for production support. Required Skills & Qualifications: - 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Platform Engineer. - Proficiency with cloud platforms such as AWS, GCP, or Azure. - Experience with containers and orchestration tools like Docker and Kubernetes. - Strong knowledge of Infrastructure as Code principles using Terraform and Ansible. - Proficiency in scripting or programming languages like Python, Go, or Bash. - Solid understanding of Linux, networking, and distributed-systems fundamentals. - Hands-on experience with CI/CD pipelines using tools like Jenkins, GitHub Actions, GitLab CI, or similar. Preferred Skills: - Experience in utilizing or building AI/ML-powered tools for operations automation, incident summaries, and alert triage. - Previous exposure to payments or fintech production environments. - Familiarity with SLO-driven reliability practices and on-call process enhancement. Techdome, a technology-driven company with over 5 years of experience in building products across various industries, offers you genuine ownership, rapid growth opportunities, and a collaborative team where your ideas are valued. The hiring process at Techdome is fast and transparent, leveraging JIA, their in-house AI hiring platform, to consistently review every application and respond within a working day. The process includes a technical round, a team discussion, and a subsequent offer. As a Site Reliability Engineer (SRE) at Techdome in Hyderabad / Indore, your role will be crucial in ensuring the availability, reliability, scalability, and performance of cloud-based production systems for our payments and platform products. You will primarily focus on automation, observability, CI/CD, incident management, and leveraging AI tooling to reduce operational toil. Key Responsibilities: - Ensure high availability, performance, and scalability of production systems. - Build automation for deployment, monitoring, and incident response. - Implement observability through metrics, logging, tracing, and alerting tools like Prometheus, Grafana, ELK, and Datadog. - Define and manage SLIs, SLOs, and error budgets. - Develop and maintain CI/CD pipelines and infrastructure as code using Terraform and Ansible. - Lead incident response, conduct root cause analysis, and participate in post-incident reviews. - Conduct capacity planning and optimize cloud costs. - Engage in an on-call rotation for production support. Required Skills & Qualifications: - 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Platform Engineer. - Proficiency with cloud platforms such as AWS, GCP, or Azure. - Experience with containers and orchestration tools like Docker and Kubernetes. - Strong knowledge of Infrastructure as Code principles using Terraform and Ansible. - Proficiency in scripting or programming languages like Python, Go, or Bash. - Solid understanding of Linux, networking, and distributed-systems fundamentals. - Hands-on experience with CI/CD pipelines using tools like Jenkins, GitHub Actions, GitLab CI, or similar. Preferred Skills: - Experience in utilizing or building AI/ML-powered tools for operations automation, incident summaries, and alert triage. - Previous exposure to payments or fintech production environments. - Familiarity with SLO-driven reliability practices and on-call process enhancement. Techdome, a technology-driven company with over 5 years of experience in building products across various industries, offers you genuine ownership, rapid growth opportunities, and a collaborative team where your ideas are valued. The hiring process at Techdome is fast and transparent, leveraging JIA, their in-house AI hiring platform, to consistently review every application and respond within a working day. The process includes a technical round, a team discussion, and a subsequent offer.

One address, no account. We’ll tell you when matching roles go live.

More at Techdome

Related open roles

View all roles