Padmi

Lead ML Ops Engineer (Uttar Pradesh)

IndiaPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Cloudkeeper

Opens the source posting on shine.com

Source description

About the role

View original

Responsibilities: - Drive R&D; and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform building the Tuner AI / Commit AI capability on GPU and ML workloads - Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement - Extend the optimization stack to LLM-era workloads caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures - Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers - Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion optimization engine customer-facing recommendation) - Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales - (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands Must Have: - B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience - Production experience with GPU workloads has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work - Strong performance engineering background must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction) - Strong Python + Linux + systems fundamentals - Solid understanding of the ML model lifecycle training, serving, inference able to reason about what is running on the GPU and why - MLOps fluency model deployment, monitoring, observability, GPU cluster operations - Hands-on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar) - Familiarity with at least one modern LLM inference framework vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML - Strong communication skills able to translate deep technical optimization into customer / business outcomes - (Lead level) Experience managing or technically leading a team of 3+ engineers Good to Have: - Deep LLM-era optimization expertise KV caching, semantic caching, model routing, energetic batching, quantization (FP16 INT8 INT4), model distillation, structured outputs - Familiarity with LLM workload patterns RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant) - CUDA, NCCL, mixed-precision training and inference - Experience with managed ML training platforms SageMaker, Azure ML, Vertex AI, Databricks Mosaic - Exposure to GPU-native clouds CoreWeave, Lambda Labs, RunPod, Crusoe - Open source contributions to ML infrastructure projects vLLM, llama.cpp, TGI, Ray, Triton, KubeRay - Adjacent experience in cloud cost optimization / FinOps Spot.io, ScaleOps, Granulate, CAST AI - Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability) Responsibilities: - Drive R&D; and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform building the Tuner AI / Commit AI capability on GPU and ML workloads - Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement - Extend the optimization stack to LLM-era workloads caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures - Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers - Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion optimization engine customer-facing recommendation) - Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales - (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands Must Have: - B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience - Production experience with GPU workloads has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work - Strong performance engineering background must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction) - Strong Python + Linux + systems fundamentals - Solid understanding of the ML model lifecycle training, serving, inference able to reason about what is running on the GPU and why - MLOps fluency model deployment, monitoring, observability, GPU cluster operations - H

One address, no account. We’ll tell you when matching roles go live.

More at Cloudkeeper

Related open roles

View all roles