Padmi

Senior Site Reliability / DevOps Engineer

ChennaiPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Datum Technologies Group

Opens the source posting on shine.com

Source description

About the role

View original

As a Site Reliability Engineer / DevOps Engineer, you will be responsible for managing and troubleshooting Linux-based systems and environments. You will develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch. Designing and managing scalable infrastructure on Microsoft Azure will also be a key part of your role. Additionally, you will administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS), and handle Kubernetes cluster lifecycle management tasks such as scaling, upgrades, troubleshooting, and maintenance. Building and maintaining CI/CD pipelines, preferably using GitHub Actions, implementing DevOps and SRE best practices, and collaborating with development teams to streamline deployment and infrastructure processes are also part of your responsibilities. Key Responsibilities: - Manage and troubleshoot Linux-based systems and environments - Develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch - Design and manage scalable infrastructure on Microsoft Azure - Administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS) - Handle Kubernetes cluster lifecycle management (scaling, upgrades, troubleshooting, and maintenance) - Build and maintain CI/CD pipelines, preferably using GitHub Actions - Implement DevOps and SRE best practices to improve system reliability and automation - Collaborate with development teams to streamline deployment and infrastructure processes Mandatory Skills: - Strong experience in Linux OS - Hands-on experience with Terraform (module development) - Experience working with Azure Cloud - Expertise in Kubernetes cluster management (AKS) - Experience with CI/CD tools (preferably GitHub Actions) Good to Have: - Experience with monitoring tools such as ELK Stack, Prometheus, or Grafana - Exposure to DevOps monitoring and automation practices We are specifically looking for candidates with hands-on experience managing and maintaining Kubernetes clusters, not just deploying applications on AKS. As a Site Reliability Engineer / DevOps Engineer, you will be responsible for managing and troubleshooting Linux-based systems and environments. You will develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch. Designing and managing scalable infrastructure on Microsoft Azure will also be a key part of your role. Additionally, you will administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS), and handle Kubernetes cluster lifecycle management tasks such as scaling, upgrades, troubleshooting, and maintenance. Building and maintaining CI/CD pipelines, preferably using GitHub Actions, implementing DevOps and SRE best practices, and collaborating with development teams to streamline deployment and infrastructure processes are also part of your responsibilities. Key Responsibilities: - Manage and troubleshoot Linux-based systems and environments - Develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch - Design and manage scalable infrastructure on Microsoft Azure - Administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS) - Handle Kubernetes cluster lifecycle management (scaling, upgrades, troubleshooting, and maintenance) - Build and maintain CI/CD pipelines, preferably using GitHub Actions - Implement DevOps and SRE best practices to improve system reliability and automation - Collaborate with development teams to streamline deployment and infrastructure processes Mandatory Skills: - Strong experience in Linux OS - Hands-on experience with Terraform (module development) - Experience working with Azure Cloud - Expertise in Kubernetes cluster management (AKS) - Experience with CI/CD tools (preferably GitHub Actions) Good to Have: - Experience with monitoring tools such as ELK Stack, Prometheus, or Grafana - Exposure to DevOps monitoring and automation practices We are specifically looking for candidates with hands-on experience managing and maintaining Kubernetes clusters, not just deploying applications on AKS.

One address, no account. We’ll tell you when matching roles go live.

More at Datum Technologies Group

Related open roles

View all roles