Source description
About the role
As an experienced AI Architect, you will be responsible for designing, building, and scaling production-ready AI voice conversation agents deployed locally and optimized for GPU-accelerated, high-throughput environments. Your role will involve owning the end-to-end architecture of real-time voice systems, including speech recognition, LLM orchestration, dialog management, speech synthesis, and low-latency streaming pipelines with a focus on reliability, scalability, and cost efficiency. This position requires a highly hands-on and strategic approach, bridging research, engineering, and production infrastructure. Key Responsibilities: - Architecture & System Design: - Design low-latency, real-time voice agent architectures for local/on-prem deployment. - Define scalable architectures for ASR LLM TTS pipelines. - Optimize systems for GPU utilization, concurrency, and throughput. - Architect fault-tolerant, production-grade voice systems (HA, monitoring, recovery). - Voice & Conversational AI: - Design and integrate ASR, LLMs, dialogue management, TTS for natural conversations. - Build streaming voice pipelines with sub-second response times. - Enable multi-turn, interruptible, natural conversations. - Model & Inference Engineering: - Deploy and optimize local LLMs and speech models. - Select and fine-tune open-source models for voice use cases. - Implement efficient inference using TensorRT, ONNX, CUDA, vLLM, Triton, or similar. - Infrastructure & Production: - Design GPU-based inference clusters. - Implement autoscaling, load balancing, and GPU scheduling. - Establish monitoring, logging, and performance metrics for voice agents. - Ensure security, privacy, and data isolation for local deployments. - Leadership & Collaboration: - Set architectural standards and best practices. - Mentor ML and platform engineers. - Collaborate with product, infra, and applied research teams. - Drive decisions from prototype to production to scale. Required Qualifications: - Technical Skills: - 7+ years in software / ML systems engineering. - 3+ years designing production AI systems. - Strong experience with real-time voice or conversational AI systems. - Hands-on experience with GPU inference optimization. - Strong Python and/or C++ background. - Experience with Linux, Docker, Kubernetes. - AI & ML Expertise: - Experience deploying open-source LLMs locally. - Knowledge of model optimization techniques. - Familiarity with voice models and systems & scaling. Preferred Qualifications: - Experience building AI voice agents or call automation systems. - Background in speech processing or audio ML. - Experience with telephony, WebRTC, SIP, or streaming audio. - Familiarity with Triton Inference Server or vLLM. - Prior experience as Tech Lead or Principal Engineer. This role offers you the opportunity to architect state-of-the-art AI voice systems, work on high-scale production deployments, competitive compensation and equity, high ownership, technical influence, and collaboration with top-tier AI and infrastructure talent. As an experienced AI Architect, you will be responsible for designing, building, and scaling production-ready AI voice conversation agents deployed locally and optimized for GPU-accelerated, high-throughput environments. Your role will involve owning the end-to-end architecture of real-time voice systems, including speech recognition, LLM orchestration, dialog management, speech synthesis, and low-latency streaming pipelines with a focus on reliability, scalability, and cost efficiency. This position requires a highly hands-on and strategic approach, bridging research, engineering, and production infrastructure. Key Responsibilities: - Architecture & System Design: - Design low-latency, real-time voice agent architectures for local/on-prem deployment. - Define scalable architectures for ASR LLM TTS pipelines. - Optimize systems for GPU utilization, concurrency, and throughput. - Architect fault-tolerant, production-grade voice systems (HA, monitoring, recovery). - Voice & Conversational AI: - Design and integrate ASR, LLMs, dialogue management, TTS for natural conversations. - Build streaming voice pipelines with sub-second response times. - Enable multi-turn, interruptible, natural conversations. - Model & Inference Engineering: - Deploy and optimize local LLMs and speech models. - Select and fine-tune open-source models for voice use cases. - Implement efficient inference using TensorRT, ONNX, CUDA, vLLM, Triton, or similar. - Infrastructure & Production: - Design GPU-based inference clusters. - Implement autoscaling, load balancing, and GPU scheduling. - Establish monitoring, logging, and performance metrics for voice agents. - Ensure security, privacy, and data isolation for local deployments. - Leadership & Collaboration: - Set architect