Padmi

DevOps (AI+MLOps) Engineer ||H || Immediate Joiner || Gurugram (Uttar Pradesh)

IndiaPosted 1 month ago
Software engineeringMid-levelFull Time; Regular
Apply at Profex Tech

Opens the source posting on shine.com

Source description

About the role

View original

Key Responsibilities: - Deploy, manage, and optimize ML/LLM models (up to 20B parameters) in production. - Design and implement scalable MLOps pipelines for training, deployment, and monitoring. - Build and optimize multi-node, multi-GPU distributed training environments. - Configure distributed training frameworks and optimize GPU utilization. - Perform model load testing, stress testing, benchmarking, and inference optimization. - Develop CI/CD pipelines for ML model deployment and retraining. - Monitor, troubleshoot, and resolve infrastructure, GPU, networking, and performance bottlenecks. - Ensure cloud-agnostic deployments across AWS, Azure, and GCP. - Collaborate with AI/ML teams to improve model reliability, scalability, and operational efficiency. Required Skills: - 4+ years of experience in DevOps/MLOps/AIOps. - Solid hands-on experience with Kubernetes, Docker, Linux, and CI/CD. - Experience with NVIDIA GPUs, CUDA, GPU optimization, and distributed training. - Hands-on experience with PyTorch and ML model deployment. - Expertise in AWS, Azure, and GCP. - Experience with Infrastructure as Code (Terraform preferred). - Knowledge of DeepSpeed, FSDP, Megatron-LM, NCCL, or similar distributed training frameworks. - Experience with vLLM, Triton Inference Server, TensorRT, or ONNX Runtime is preferred. - Strong scripting skills in Python and Shell. - Experience with monitoring tools such as Prometheus, Grafana, or OpenTelemetry. Preferred Skills: - Experience deploying and managing Large Language Models (LLMs). - Knowledge of Hugging Face ecosystem and Generative AI. - Understanding of model serving, inference optimization, autoscaling, and cost optimization. - Exposure to ML observability and AI infrastructure. .

One address, no account. We’ll tell you when matching roles go live.

More at Profex Tech

Related open roles

View all roles