Source description
About the role
Role Overview: As an AI Inference Engineer at Qubrid AI, your main responsibility will be to deploy, optimize, and operate open-source and commercial AI models across NVIDIA GPU infrastructure. You will work at the intersection of machine learning, distributed systems, GPU optimization, and cloud infrastructure to deliver low-latency, high-throughput AI services. This role requires deep expertise in LLM serving, GPU performance tuning, model optimization, inference frameworks, and large-scale production deployments. Responsibilities: - Deploy and manage Large Language Models (LLMs), multimodal models, vision models, speech models, and embedding models in production. - Build and optimize inference pipelines for enterprise and public AI workloads. - Implement scalable serving architectures using modern inference frameworks. - Support model versioning, rollbacks, canary deployments, and A/B testing. - Optimize GPU utilization, memory allocation, throughput, and latency. - Implement model quantization techniques including FP16, BF16, INT8, GPTQ, AWQ, and GGUF. - Tune inference workloads across NVIDIA H100, H200, B300, B200, A100, L40S, and other accelerator platforms. - Design scalable inference clusters using Kubernetes and containerized workloads. - Implement auto-scaling, load balancing, and fault-tolerant architectures. - Build GPU scheduling and resource allocation strategies. - Deploy and optimize models using various frameworks like vLLM, NVIDIA TensorRT-LLM, Triton Inference Server, SGLang, TGI, Ollama, Ray Serve, and more. - Implement batching, continuous batching, speculative decoding, KV cache optimization, and context caching. - Develop APIs and backend services supporting AI inference workloads. - Integrate authentication, billing, token metering, and usage tracking. - Contribute to Qubrid's AI Model Studio and AI Compute Platform. Qualification Required: - Bachelor's or Master's degree in Computer Science, Engineering, AI/ML, or related field. - 2+ years of software engineering experience. - 2+ years of production AI/ML infrastructure experience. - Strong Python programming expertise. - Deep understanding of transformer architectures and modern LLMs. - Experience with Linux systems administration, Docker, Kubernetes, and distributed systems. - Familiarity with AI & ML tools like PyTorch, Hugging Face Transformers, model quantization, fine-tuning workflows, and more. - Knowledge of GPU & Infrastructure technologies such as NVIDIA CUDA, TensorRT, NCCL, NVLink, and multi-GPU optimization. - Experience with Cloud & DevOps tools like Kubernetes, Docker, Terraform, CI/CD pipelines, and cloud environments. - Proficiency in Databases & Backend technologies including PostgreSQL, MongoDB, Redis, REST APIs, gRPC, and event-driven architectures. Role Overview: As an AI Inference Engineer at Qubrid AI, your main responsibility will be to deploy, optimize, and operate open-source and commercial AI models across NVIDIA GPU infrastructure. You will work at the intersection of machine learning, distributed systems, GPU optimization, and cloud infrastructure to deliver low-latency, high-throughput AI services. This role requires deep expertise in LLM serving, GPU performance tuning, model optimization, inference frameworks, and large-scale production deployments. Responsibilities: - Deploy and manage Large Language Models (LLMs), multimodal models, vision models, speech models, and embedding models in production. - Build and optimize inference pipelines for enterprise and public AI workloads. - Implement scalable serving architectures using modern inference frameworks. - Support model versioning, rollbacks, canary deployments, and A/B testing. - Optimize GPU utilization, memory allocation, throughput, and latency. - Implement model quantization techniques including FP16, BF16, INT8, GPTQ, AWQ, and GGUF. - Tune inference workloads across NVIDIA H100, H200, B300, B200, A100, L40S, and other accelerator platforms. - Design scalable inference clusters using Kubernetes and containerized workloads. - Implement auto-scaling, load balancing, and fault-tolerant architectures. - Build GPU scheduling and resource allocation strategies. - Deploy and optimize models using various frameworks like vLLM, NVIDIA TensorRT-LLM, Triton Inference Server, SGLang, TGI, Ollama, Ray Serve, and more. - Implement batching, continuous batching, speculative decoding, KV cache optimization, and context caching. - Develop APIs and backend services supporting AI inference workloads. - Integrate authentication, billing, token metering, and usage tracking. - Contribute to Qubrid's AI Model Studio and AI Compute Platform. Qualification Required: - Bachelor's or Master's degree in Computer Science, Engineering, AI/ML, or related field. - 2+ years of software engineering experience. - 2+ years of production AI/ML infrastructure experience. - Strong Python programming expertise. - Deep understanding of transformer architectures and modern LLMs. - E
More at Qubrid AI
Related open roles
GPU Kubernetes Cluster Engineer (Junior) WFH
India
Junior Back End Engineer Ai Wfh Kolkata (India)
India
Junior Network Engineer SONiC Open Networking Software WFH (Delhi)
Delhi NCR
Junior Network Engineer SONiC Open Networking Software WFH (Coimbatore)
India
Junior Software Engineer, AI Platforms (Kolkata)
India
Enterprise Software Engineer (Junior)
India