Source description
About the role
In this role, you will be responsible for designing, building, and owning scalable, secure, and highly available cloud infrastructure. You will also be in charge of the end-to-end ownership of production environments, including uptime, performance, and cost optimization. Your key responsibilities will include building and maintaining CI/CD pipelines for fast and reliable releases, implementing and managing infrastructure as code using tools like Terraform, Pulumi, or CloudFormation, leading incident response and post-mortems, and establishing observability systems. You will partner closely with engineering and ML teams to support AI workloads and data pipelines, drive security, access control, and compliance best practices, mentor junior engineers, and deploy on private cloud environments. Your day-to-day activities will involve reviewing and improving cloud architecture and deployment strategies, debugging and resolving real production issues, automating infrastructure and operational workflows, making trade-offs between speed, cost, and reliability, and setting DevOps standards to enhance the overall engineering practices. Qualifications Required: - 10 years+ of hands-on experience in DevOps, Platform Engineering, or SRE roles - Tier-1 engineering pedigree (IITs / IISc / top global CS programs) - Strong experience with at least one major cloud provider (GCP and AWS) - Strong fundamentals in Linux, networking, and distributed systems - Production experience with Docker and Kubernetes - Experience building and maintaining CI/CD pipelines - Hands-on experience with infrastructure as code tools - Proven experience owning and supporting production systems In this role, you will be responsible for designing, building, and owning scalable, secure, and highly available cloud infrastructure. You will also be in charge of the end-to-end ownership of production environments, including uptime, performance, and cost optimization. Your key responsibilities will include building and maintaining CI/CD pipelines for fast and reliable releases, implementing and managing infrastructure as code using tools like Terraform, Pulumi, or CloudFormation, leading incident response and post-mortems, and establishing observability systems. You will partner closely with engineering and ML teams to support AI workloads and data pipelines, drive security, access control, and compliance best practices, mentor junior engineers, and deploy on private cloud environments. Your day-to-day activities will involve reviewing and improving cloud architecture and deployment strategies, debugging and resolving real production issues, automating infrastructure and operational workflows, making trade-offs between speed, cost, and reliability, and setting DevOps standards to enhance the overall engineering practices. Qualifications Required: - 10 years+ of hands-on experience in DevOps, Platform Engineering, or SRE roles - Tier-1 engineering pedigree (IITs / IISc / top global CS programs) - Strong experience with at least one major cloud provider (GCP and AWS) - Strong fundamentals in Linux, networking, and distributed systems - Production experience with Docker and Kubernetes - Experience building and maintaining CI/CD pipelines - Hands-on experience with infrastructure as code tools - Proven experience owning and supporting production systems