Source description
About the role
Bachelor's or Master's in Computer Science or a related field.
9+ years of software engineering experience, including 2+ years managing engineers.
Strong hands-on experience with the modern LLM inference stack — TensorRT-LLM, vLLM, SGLang — and with production, low-latency model serving at scale.
Depth in inference optimization: GPU kernel tuning, quantization, speculative decoding, KV-cache and memory/IO optimization. CUDA / GPU programming experience is a strong plus.
Experience with distributed training and the frameworks behind it — PyTorch FSDP, DeepSpeed, Megatron, or Ray.
Experience running GPU fleets in production — Kubernetes (ideally GKE), GPU scheduling and allocation, and multi-region/multi-cluster deployment.
Familiarity with building LLM-powered agents and agentic workflows, and a point of view on where autonomy can replace manual engineering toil.
Experience with big-data and streaming stacks — Spark, Flink, or similar.
Proficiency in Python; systems-level fluency (C++ / Go / Rust) for performance-critical paths.
Strong leadership, problem-solving, and stakeholder-management skills.
Preferred
Open-source contributions to inference engines, training frameworks, or ML infra tooling.
Experience managing GPU cost/efficiency (FinOps) for a large fleet on Cloud and Neo-Clouds.
Track record building platforms for high-scale consumer products (millions of users).
Familiarity with observability and reliability for ML systems (SLOs, autoscaling, incident response).
More at Meesho
Related open roles
Engineering Manager- Database Platform
Bangalore · Onsite
Engineering Manager – AI Engineering
India
Engineering Manager- Infrastructure platform
Bangalore
Engineering Manager- Data Platform
Bangalore
Engineering Manager- Infrastructure platform
Bangalore · Onsite
Engineering Manager- Data Platform
Bangalore · Onsite
