Source description
About the role
As a Site Reliability Engineer / DevOps Engineer, you will be responsible for managing and troubleshooting Linux-based systems and environments. You will develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch. Designing and managing scalable infrastructure on Microsoft Azure will also be a key part of your role. Additionally, you will administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS), and handle Kubernetes cluster lifecycle management tasks such as scaling, upgrades, troubleshooting, and maintenance. Building and maintaining CI/CD pipelines, preferably using GitHub Actions, implementing DevOps and SRE best practices, and collaborating with development teams to streamline deployment and infrastructure processes are also part of your responsibilities. Key Responsibilities: - Manage and troubleshoot Linux-based systems and environments - Develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch - Design and manage scalable infrastructure on Microsoft Azure - Administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS) - Handle Kubernetes cluster lifecycle management (scaling, upgrades, troubleshooting, and maintenance) - Build and maintain CI/CD pipelines, preferably using GitHub Actions - Implement DevOps and SRE best practices to improve system reliability and automation - Collaborate with development teams to streamline deployment and infrastructure processes Mandatory Skills: - Strong experience in Linux OS - Hands-on experience with Terraform (module development) - Experience working with Azure Cloud - Expertise in Kubernetes cluster management (AKS) - Experience with CI/CD tools (preferably GitHub Actions) Good to Have: - Experience with monitoring tools such as ELK Stack, Prometheus, or Grafana - Exposure to DevOps monitoring and automation practices We are specifically looking for candidates with hands-on experience managing and maintaining Kubernetes clusters, not just deploying applications on AKS. As a Site Reliability Engineer / DevOps Engineer, you will be responsible for managing and troubleshooting Linux-based systems and environments. You will develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch. Designing and managing scalable infrastructure on Microsoft Azure will also be a key part of your role. Additionally, you will administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS), and handle Kubernetes cluster lifecycle management tasks such as scaling, upgrades, troubleshooting, and maintenance. Building and maintaining CI/CD pipelines, preferably using GitHub Actions, implementing DevOps and SRE best practices, and collaborating with development teams to streamline deployment and infrastructure processes are also part of your responsibilities. Key Responsibilities: - Manage and troubleshoot Linux-based systems and environments - Develop and maintain Infrastructure as Code using Terraform, including writing Terraform modules from scratch - Design and manage scalable infrastructure on Microsoft Azure - Administer and maintain Kubernetes clusters, particularly Azure Kubernetes Service (AKS) - Handle Kubernetes cluster lifecycle management (scaling, upgrades, troubleshooting, and maintenance) - Build and maintain CI/CD pipelines, preferably using GitHub Actions - Implement DevOps and SRE best practices to improve system reliability and automation - Collaborate with development teams to streamline deployment and infrastructure processes Mandatory Skills: - Strong experience in Linux OS - Hands-on experience with Terraform (module development) - Experience working with Azure Cloud - Expertise in Kubernetes cluster management (AKS) - Experience with CI/CD tools (preferably GitHub Actions) Good to Have: - Experience with monitoring tools such as ELK Stack, Prometheus, or Grafana - Exposure to DevOps monitoring and automation practices We are specifically looking for candidates with hands-on experience managing and maintaining Kubernetes clusters, not just deploying applications on AKS.
More at Datum Technologies Group