Source description
About the role
Role Overview: As a Site Reliability Engineer (SRE) in Platform Engineering with experience in Gen AI, you will be responsible for building, scaling, and operating reliable, secure, and automated platform services. Your main focus will be on improving system reliability, reducing operational toil, and enabling developer productivity through strong engineering practices. Additionally, you will have the opportunity to work on COE initiatives and internal Intellectual properties and products to meet various client requirements. Key Responsibilities: - Build and operate highly available, scalable platform services. - Implement SRE practices such as SLIs, SLOs, error budgets, and automation. - Manage and scale Kubernetes-based workloads in cloud environments. - Develop and maintain CI/CD pipelines and Infrastructure as Code. - Own production reliability, participate in on-call rotations, incident response, and RCA (if required). - Enhance observability using metrics, logs, and traces. Qualifications Required: - 5-16 years of experience in SRE, Platform, DevOps, Cloud Engineering, or AI. - Hands-on experience with Kubernetes, Docker, and Linux. - Strong scripting/programming skills in Python, Go, and Bash. - Experience with AWS, Azure, or GCP. - Familiarity with Terraform, Infrastructure as Code (IaC), monitoring, and alerting tools. Additional Company Details: - Nice to have experience with internal developer platforms, service mesh, or FinOps. - Cloud or Kubernetes certifications are a plus. Role Overview: As a Site Reliability Engineer (SRE) in Platform Engineering with experience in Gen AI, you will be responsible for building, scaling, and operating reliable, secure, and automated platform services. Your main focus will be on improving system reliability, reducing operational toil, and enabling developer productivity through strong engineering practices. Additionally, you will have the opportunity to work on COE initiatives and internal Intellectual properties and products to meet various client requirements. Key Responsibilities: - Build and operate highly available, scalable platform services. - Implement SRE practices such as SLIs, SLOs, error budgets, and automation. - Manage and scale Kubernetes-based workloads in cloud environments. - Develop and maintain CI/CD pipelines and Infrastructure as Code. - Own production reliability, participate in on-call rotations, incident response, and RCA (if required). - Enhance observability using metrics, logs, and traces. Qualifications Required: - 5-16 years of experience in SRE, Platform, DevOps, Cloud Engineering, or AI. - Hands-on experience with Kubernetes, Docker, and Linux. - Strong scripting/programming skills in Python, Go, and Bash. - Experience with AWS, Azure, or GCP. - Familiarity with Terraform, Infrastructure as Code (IaC), monitoring, and alerting tools. Additional Company Details: - Nice to have experience with internal developer platforms, service mesh, or FinOps. - Cloud or Kubernetes certifications are a plus.
More at LTM