Source description
About the role
We are looking for an experienced Inference Engineer to build and optimize high-performance inference systems for Large Language Models (LLMs), Speech AI systems, and multimodal AI workloads. You will work on deploying production-grade AI systems with a strong focus on: low latency, high throughput, GPU efficiency, scalable serving infrastructure, distributed inference, and cost optimization. This role sits at the intersection of: systems engineering, deep learning infrastructure, distributed computing, and production AI deployment. You will collaborate closely with: ML researchers, platform engineers, speech AI teams, and product engineering teams.
More at Soket AI