Padmi

AI Architect

HyderabadPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at VMax eSolutions India Pvt Ltd

Opens the source posting on shine.com

Source description

About the role

View original

As an experienced AI Architect, you will be responsible for designing, building, and scaling production-ready AI voice conversation agents deployed locally and optimized for GPU-accelerated, high-throughput environments. Your role will involve owning the end-to-end architecture of real-time voice systems, including speech recognition, LLM orchestration, dialog management, speech synthesis, and low-latency streaming pipelines with a focus on reliability, scalability, and cost efficiency. This position requires a highly hands-on and strategic approach, bridging research, engineering, and production infrastructure. Key Responsibilities: - Architecture & System Design: - Design low-latency, real-time voice agent architectures for local/on-prem deployment. - Define scalable architectures for ASR LLM TTS pipelines. - Optimize systems for GPU utilization, concurrency, and throughput. - Architect fault-tolerant, production-grade voice systems (HA, monitoring, recovery). - Voice & Conversational AI: - Design and integrate ASR, LLMs, dialogue management, TTS for natural conversations. - Build streaming voice pipelines with sub-second response times. - Enable multi-turn, interruptible, natural conversations. - Model & Inference Engineering: - Deploy and optimize local LLMs and speech models. - Select and fine-tune open-source models for voice use cases. - Implement efficient inference using TensorRT, ONNX, CUDA, vLLM, Triton, or similar. - Infrastructure & Production: - Design GPU-based inference clusters. - Implement autoscaling, load balancing, and GPU scheduling. - Establish monitoring, logging, and performance metrics for voice agents. - Ensure security, privacy, and data isolation for local deployments. - Leadership & Collaboration: - Set architectural standards and best practices. - Mentor ML and platform engineers. - Collaborate with product, infra, and applied research teams. - Drive decisions from prototype to production to scale. Required Qualifications: - Technical Skills: - 7+ years in software / ML systems engineering. - 3+ years designing production AI systems. - Strong experience with real-time voice or conversational AI systems. - Hands-on experience with GPU inference optimization. - Strong Python and/or C++ background. - Experience with Linux, Docker, Kubernetes. - AI & ML Expertise: - Experience deploying open-source LLMs locally. - Knowledge of model optimization techniques. - Familiarity with voice models and systems & scaling. Preferred Qualifications: - Experience building AI voice agents or call automation systems. - Background in speech processing or audio ML. - Experience with telephony, WebRTC, SIP, or streaming audio. - Familiarity with Triton Inference Server or vLLM. - Prior experience as Tech Lead or Principal Engineer. This role offers you the opportunity to architect state-of-the-art AI voice systems, work on high-scale production deployments, competitive compensation and equity, high ownership, technical influence, and collaboration with top-tier AI and infrastructure talent. As an experienced AI Architect, you will be responsible for designing, building, and scaling production-ready AI voice conversation agents deployed locally and optimized for GPU-accelerated, high-throughput environments. Your role will involve owning the end-to-end architecture of real-time voice systems, including speech recognition, LLM orchestration, dialog management, speech synthesis, and low-latency streaming pipelines with a focus on reliability, scalability, and cost efficiency. This position requires a highly hands-on and strategic approach, bridging research, engineering, and production infrastructure. Key Responsibilities: - Architecture & System Design: - Design low-latency, real-time voice agent architectures for local/on-prem deployment. - Define scalable architectures for ASR LLM TTS pipelines. - Optimize systems for GPU utilization, concurrency, and throughput. - Architect fault-tolerant, production-grade voice systems (HA, monitoring, recovery). - Voice & Conversational AI: - Design and integrate ASR, LLMs, dialogue management, TTS for natural conversations. - Build streaming voice pipelines with sub-second response times. - Enable multi-turn, interruptible, natural conversations. - Model & Inference Engineering: - Deploy and optimize local LLMs and speech models. - Select and fine-tune open-source models for voice use cases. - Implement efficient inference using TensorRT, ONNX, CUDA, vLLM, Triton, or similar. - Infrastructure & Production: - Design GPU-based inference clusters. - Implement autoscaling, load balancing, and GPU scheduling. - Establish monitoring, logging, and performance metrics for voice agents. - Ensure security, privacy, and data isolation for local deployments. - Leadership & Collaboration: - Set architect

One address, no account. We’ll tell you when matching roles go live.