Padmi

SRE Engineer

Delhi NCRPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Proximity Labs

Opens the source posting on shine.com

Source description

About the role

View original

Responsibilities: Own day-2 production operations of a large-scale, AI-first platform running on cloud infrastructure.Run, scale, and harden Kubernetes-based workloads integrated with a broad set of managed cloud services across data, messaging, AI, networking, and security.Define, implement, and operate SLIs, SLOs, and error budgets across core platform and AI services.Build and own observability end-to-end, including: APM, Infrastructure monitoring, Logs, alerts, and operational dashboards.Improve and maintain CI/CD pipelines and Terraform-driven infrastructure automation.Operate and integrate AI platform services for LLM deployments and model lifecycle management.Lead incident response, conduct blameless postmortems, and drive systemic reliability improvements.Optimise cost, performance, and autoscaling for AI, ML, and data-intensive workloads.Partner closely with backend, data, and ML engineers to ensure production readiness and operational best practices. What Matters (Non-Negotiable Alignment) Infra owners, not operators. This role is for engineers who design, build, and own infrastructure, not those limited to ticket-based operations. Built and operated production-grade cloud infrastructure end-to-end.Strong Kubernetes experience in real, high-traffic production environments.AWS experience is mandatory, with GCP as a strong plus.Experience operating AI / ML workloads in production.Including GPU-based systems.Strong ownership of CI/CD systems and Infrastructure as Code.End-to-end observability ownership.Monitoring, logging, alerting, dashboards.Comfortable making infrastructure decisions under ambiguity.Proven ability to collaborate deeply with ML and backend teams to take systems from design to production scale. Requirements: 6+ years of hands-on experience in DevOps, SRE, or Platform Engineering roles.Strong, production-grade experience with cloud platforms.AWS required.GCP strongly preferred, especially Kubernetes and managed services.Proven expertise running Kubernetes at scale in live production environments.Deep hands-on experience with New Relic in complex, distributed systems.Experience operating AI/ML or LLM-driven platforms in production environments.Solid background in Terraform, CI/CD systems, cloud networking, and security fundamentals.Strong understanding of reliability engineering principles, including capacity planning, failure modes, and resilience patterns.Comfortable owning production systems end-to-end with minimal supervision.Strong communication skills and the ability to operate calmly and effectively during incidents.Experience building internal platform tooling for developer productivity. Desired Skills: Experience managing multi-cloud environments or cross-cloud integrations.Familiarity with cost optimisation strategies for large-scale Kubernetes and AI workloads.Exposure to service meshes, advanced traffic management, or zero-trust security models. Responsibilities: Own day-2 production operations of a large-scale, AI-first platform running on cloud infrastructure.Run, scale, and harden Kubernetes-based workloads integrated with a broad set of managed cloud services across data, messaging, AI, networking, and security.Define, implement, and operate SLIs, SLOs, and error budgets across core platform and AI services.Build and own observability end-to-end, including: APM, Infrastructure monitoring, Logs, alerts, and operational dashboards.Improve and maintain CI/CD pipelines and Terraform-driven infrastructure automation.Operate and integrate AI platform services for LLM deployments and model lifecycle management.Lead incident response, conduct blameless postmortems, and drive systemic reliability improvements.Optimise cost, performance, and autoscaling for AI, ML, and data-intensive workloads.Partner closely with backend, data, and ML engineers to ensure production readiness and operational best practices. What Matters (Non-Negotiable Alignment) Infra owners, not operators. This role is for engineers who design, build, and own infrastructure, not those limited to ticket-based operations. Built and operated production-grade cloud infrastructure end-to-end.Strong Kubernetes experience in real, high-traffic production environments.AWS experience is mandatory, with GCP as a strong plus.Experience operating AI / ML workloads in production.Including GPU-based systems.Strong ownership of CI/CD systems and Infrastructure as Code.End-to-end observability ownership.Monitoring, logging, alerting, dashboards.Comfortable making infrastructure decisions under ambiguity.Proven ability to collaborate deeply with ML and backend teams to take systems from design to production scale. Requirements: 6+ years of hands-on experience in DevOps, SRE, or Platform Engineering roles.Strong, production-grade experience with cloud platforms.AWS required.GCP strongly preferred, especially Kubernetes and managed services.Proven expertise running Kubernetes at scale in live production

One address, no account. We’ll tell you when matching roles go live.

More at Proximity Labs

Related open roles

View all roles