Padmi

DevOps (MLOps+AIOps) Engineer

Delhi NCRPosted 1 month ago
Software engineeringMid-level
Apply at Profex Tech

Opens the source posting on naukri.com

Source description

About the role

View original

Key Roles & Responsibilities • Deploy and manage ML/LLM models in production environments. • Design and execute model load testing and performance benchmarking (latency, throughput, memory, cost). • Build and optimize multi-node, multi-GPU training pipelines. • Configure and tune distributed training frameworks (data parallelism, model parallelism, pipeline parallelism). • Optimize GPU utilization, memory footprint, and inference costs. • Set up CI/CD pipelines for model deployment and retraining. • Troubleshoot GPU, networking, and performance bottlenecks. • Work across cloud platforms to ensure portability and vendor-agnostic deployments. Required Skill Sets • Strong experience with GPU workloads (NVIDIA GPUs, CUDA concepts). • Proven expertise in model deployment on AWS, Azure, and GCP. • Hands-on experience deploying models up to 20B parameters. • Experience with distributed training (multi-node, multi-GPU setups). • Deep understanding of load testing, stress testing, and benchmarking ML systems.

One address, no account. We’ll tell you when matching roles go live.

More at Profex Tech

Related open roles

View all roles