Padmi

DevOps SRE + Gen AI

MumbaiPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at LTM

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: As a Site Reliability Engineer (SRE) in Platform Engineering with experience in Gen AI, you will be responsible for building, scaling, and operating reliable, secure, and automated platform services. Your main focus will be on improving system reliability, reducing operational toil, and enabling developer productivity through strong engineering practices. Additionally, you will have the opportunity to work on COE initiatives and internal Intellectual properties and products to meet various client requirements. Key Responsibilities: - Build and operate highly available, scalable platform services. - Implement SRE practices such as SLIs, SLOs, error budgets, and automation. - Manage and scale Kubernetes-based workloads in cloud environments. - Develop and maintain CI/CD pipelines and Infrastructure as Code. - Own production reliability, participate in on-call rotations, incident response, and RCA (if required). - Enhance observability using metrics, logs, and traces. Qualifications Required: - 5-16 years of experience in SRE, Platform, DevOps, Cloud Engineering, or AI. - Hands-on experience with Kubernetes, Docker, and Linux. - Strong scripting/programming skills in Python, Go, and Bash. - Experience with AWS, Azure, or GCP. - Familiarity with Terraform, Infrastructure as Code (IaC), monitoring, and alerting tools. Additional Company Details: - Nice to have experience with internal developer platforms, service mesh, or FinOps. - Cloud or Kubernetes certifications are a plus. Role Overview: As a Site Reliability Engineer (SRE) in Platform Engineering with experience in Gen AI, you will be responsible for building, scaling, and operating reliable, secure, and automated platform services. Your main focus will be on improving system reliability, reducing operational toil, and enabling developer productivity through strong engineering practices. Additionally, you will have the opportunity to work on COE initiatives and internal Intellectual properties and products to meet various client requirements. Key Responsibilities: - Build and operate highly available, scalable platform services. - Implement SRE practices such as SLIs, SLOs, error budgets, and automation. - Manage and scale Kubernetes-based workloads in cloud environments. - Develop and maintain CI/CD pipelines and Infrastructure as Code. - Own production reliability, participate in on-call rotations, incident response, and RCA (if required). - Enhance observability using metrics, logs, and traces. Qualifications Required: - 5-16 years of experience in SRE, Platform, DevOps, Cloud Engineering, or AI. - Hands-on experience with Kubernetes, Docker, and Linux. - Strong scripting/programming skills in Python, Go, and Bash. - Experience with AWS, Azure, or GCP. - Familiarity with Terraform, Infrastructure as Code (IaC), monitoring, and alerting tools. Additional Company Details: - Nice to have experience with internal developer platforms, service mesh, or FinOps. - Cloud or Kubernetes certifications are a plus.

One address, no account. We’ll tell you when matching roles go live.

More at LTM

Related open roles

View all roles