Padmi
Nava logo
Nava

government technology modernization · benefits administration platforms

Post-Training Optimization Engineer (LLM Inference & Efficiency)

BangalorePosted 2 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Nava

Opens the source posting on shine.com

Source description

About the role

View original

Role & Responsibilities Optimize LLM inference pipelines for latency, throughput, and memory efficiency across GPU/TPU hardwareusing quantization, pruning, kernel fusion, and runtime scheduling.Implement and benchmark model compression techniques (INT4/INT8/FP8, LoRA, Sparse attention) to reduce model size while preserving performance.Integrate and tune LLM serving frameworks (vLLM, TensorRT-LLM, HuggingFace TGI, ONNX Runtime) for high-throughput, low-latency inference in cloud and on-premise environments.Collaborate with ML researchers and infrastructure engineers to profile bottlenecks and implement hardware-aware optimizations using CUDA, Triton, or custom kernels.Design and automate model benchmarking suites to track performance regression, memory footprint, and cost-per-token across model versions and hardware targets.Document optimization playbooks and contribute to internal tooling for repeatable, scalable model deployment workflows. Skills & Qualifications Must-Have PyTorchTensorRTvLLMQuantization (INT4/INT8/FP8)CUDAONNX RuntimeTriton Inference ServerLLM Inference Optimization Preferred Experience with Mixture-of-Experts (MoE) modelsFamiliarity with HuggingFace Transformers and TGIKnowledge of NPU/ASIC inference backends (e.g., Qualcomm, Groq, Cerebras) Skills: cuda,ml,learning,compression,optimization,distillation,decoding,research,machine learning,training Role & Responsibilities Optimize LLM inference pipelines for latency, throughput, and memory efficiency across GPU/TPU hardwareusing quantization, pruning, kernel fusion, and runtime scheduling.Implement and benchmark model compression techniques (INT4/INT8/FP8, LoRA, Sparse attention) to reduce model size while preserving performance.Integrate and tune LLM serving frameworks (vLLM, TensorRT-LLM, HuggingFace TGI, ONNX Runtime) for high-throughput, low-latency inference in cloud and on-premise environments.Collaborate with ML researchers and infrastructure engineers to profile bottlenecks and implement hardware-aware optimizations using CUDA, Triton, or custom kernels.Design and automate model benchmarking suites to track performance regression, memory footprint, and cost-per-token across model versions and hardware targets.Document optimization playbooks and contribute to internal tooling for repeatable, scalable model deployment workflows. Skills & Qualifications Must-Have PyTorchTensorRTvLLMQuantization (INT4/INT8/FP8)CUDAONNX RuntimeTriton Inference ServerLLM Inference Optimization Preferred Experience with Mixture-of-Experts (MoE) modelsFamiliarity with HuggingFace Transformers and TGIKnowledge of NPU/ASIC inference backends (e.g., Qualcomm, Groq, Cerebras) Skills: cuda,ml,learning,compression,optimization,distillation,decoding,research,machine learning,training

One address, no account. We’ll tell you when matching roles go live.

More at Nava

Related open roles

View all roles