Padmi

AI Inference Junior Engineer WFH

IndiaPosted 2 months ago
Software engineeringJuniorFull Time; Regular
Apply at Qubrid AI

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: As an AI Inference Engineer at Qubrid AI, your main responsibility will be to deploy, optimize, and operate open-source and commercial AI models across NVIDIA GPU infrastructure. You will work at the intersection of machine learning, distributed systems, GPU optimization, and cloud infrastructure to deliver low-latency, high-throughput AI services. This role requires deep expertise in LLM serving, GPU performance tuning, model optimization, inference frameworks, and large-scale production deployments. Responsibilities: - Deploy and manage Large Language Models (LLMs), multimodal models, vision models, speech models, and embedding models in production. - Build and optimize inference pipelines for enterprise and public AI workloads. - Implement scalable serving architectures using modern inference frameworks. - Support model versioning, rollbacks, canary deployments, and A/B testing. - Optimize GPU utilization, memory allocation, throughput, and latency. - Implement model quantization techniques including FP16, BF16, INT8, GPTQ, AWQ, and GGUF. - Tune inference workloads across NVIDIA H100, H200, B300, B200, A100, L40S, and other accelerator platforms. - Design scalable inference clusters using Kubernetes and containerized workloads. - Implement auto-scaling, load balancing, and fault-tolerant architectures. - Build GPU scheduling and resource allocation strategies. - Deploy and optimize models using various frameworks like vLLM, NVIDIA TensorRT-LLM, Triton Inference Server, SGLang, TGI, Ollama, Ray Serve, and more. - Implement batching, continuous batching, speculative decoding, KV cache optimization, and context caching. - Develop APIs and backend services supporting AI inference workloads. - Integrate authentication, billing, token metering, and usage tracking. - Contribute to Qubrid's AI Model Studio and AI Compute Platform. Qualification Required: - Bachelor's or Master's degree in Computer Science, Engineering, AI/ML, or related field. - 2+ years of software engineering experience. - 2+ years of production AI/ML infrastructure experience. - Strong Python programming expertise. - Deep understanding of transformer architectures and modern LLMs. - Experience with Linux systems administration, Docker, Kubernetes, and distributed systems. - Familiarity with AI & ML tools like PyTorch, Hugging Face Transformers, model quantization, fine-tuning workflows, and more. - Knowledge of GPU & Infrastructure technologies such as NVIDIA CUDA, TensorRT, NCCL, NVLink, and multi-GPU optimization. - Experience with Cloud & DevOps tools like Kubernetes, Docker, Terraform, CI/CD pipelines, and cloud environments. - Proficiency in Databases & Backend technologies including PostgreSQL, MongoDB, Redis, REST APIs, gRPC, and event-driven architectures. Role Overview: As an AI Inference Engineer at Qubrid AI, your main responsibility will be to deploy, optimize, and operate open-source and commercial AI models across NVIDIA GPU infrastructure. You will work at the intersection of machine learning, distributed systems, GPU optimization, and cloud infrastructure to deliver low-latency, high-throughput AI services. This role requires deep expertise in LLM serving, GPU performance tuning, model optimization, inference frameworks, and large-scale production deployments. Responsibilities: - Deploy and manage Large Language Models (LLMs), multimodal models, vision models, speech models, and embedding models in production. - Build and optimize inference pipelines for enterprise and public AI workloads. - Implement scalable serving architectures using modern inference frameworks. - Support model versioning, rollbacks, canary deployments, and A/B testing. - Optimize GPU utilization, memory allocation, throughput, and latency. - Implement model quantization techniques including FP16, BF16, INT8, GPTQ, AWQ, and GGUF. - Tune inference workloads across NVIDIA H100, H200, B300, B200, A100, L40S, and other accelerator platforms. - Design scalable inference clusters using Kubernetes and containerized workloads. - Implement auto-scaling, load balancing, and fault-tolerant architectures. - Build GPU scheduling and resource allocation strategies. - Deploy and optimize models using various frameworks like vLLM, NVIDIA TensorRT-LLM, Triton Inference Server, SGLang, TGI, Ollama, Ray Serve, and more. - Implement batching, continuous batching, speculative decoding, KV cache optimization, and context caching. - Develop APIs and backend services supporting AI inference workloads. - Integrate authentication, billing, token metering, and usage tracking. - Contribute to Qubrid's AI Model Studio and AI Compute Platform. Qualification Required: - Bachelor's or Master's degree in Computer Science, Engineering, AI/ML, or related field. - 2+ years of software engineering experience. - 2+ years of production AI/ML infrastructure experience. - Strong Python programming expertise. - Deep understanding of transformer architectures and modern LLMs. - E

One address, no account. We’ll tell you when matching roles go live.

More at Qubrid AI

Related open roles

View all roles
AI Inference Junior Engineer WFH at Qubrid AI · Padmi