Source description
About the role
Responsibilities Senior Staff Engineer, Automatic Speech Recognition (ASR) Location: Bengaluru Type: Full-time Function: AI Team About the role While youre reading this, Gnani is listening to thousands of customers across India in Hindi, Tamil, Telugu, and dozens more languages, over noisy phone lines, thick accents, and mid-sentence code-switches. That understanding is our ASR stack. Its the difference between a voice agent that mishears and one that gets it right the first time, in the language and accent the caller actually uses. Were building the most accurate speech recognition for Indian languages, in real time and at enterprise scale and we want you to lead the research behind it. As Senior Staff Engineer for ASR, you own the stack end-to-end: advancing accuracy and multilingual coverage, engineering it for streaming latency, and hardening it for the messy realities of production audio. This is a senior individual-contributor role: you set technical direction, design experiments, and mentor the engineers around you without a people-management load. Core mandate Research Advance ASR research for Indian languages accuracy, multilingual coverage, and robustness across modern architectures Production performance Ship streaming and offline ASR with low latency, high throughput, and reliability at scale Voice AI ASR Feature Enhancements Handle real-world audio and build the real-time components conversations depend on Technical leadership Guide engineers, design rigorous experiments, and own the standard for how ASR gets built here What youll drive Research modeling - Own the ASR modeling roadmap architecture selection, accuracy, and multilingual / code-switched coverage across Indian languages - Work across modern ASR architectures (Conformer, E-Branchformer, RNN-T, CTC, Whisper / encoder-decoder, and Speech Language Models) and translate findings into shippable systems - Push on the hard multilingual problems: code-switching, regional accents, dialects, pronunciation variation, and low-resource languages - Train and fine-tune large-scale multilingual models data collection, augmentation, domain adaptation, contextual biasing, and custom vocabulary - Own the text side of recognition: tokenization (BPE, SentencePiece, phonemes), language modeling, confidence scoring, punctuation, capitalization, and text normalization - Participate in benchmarks and publish the work in international conferences Production performance engineering - Design ASR systems for production streaming (low partial / first-token latency) and offline (high-throughput batch), with favourable RTF - Optimize production inference using ONNX, TensorRT, NVIDIA Triton, quantization, batching, and GPU-productive serving architectures - Partner with platform and infra teams on serving, scaling, and reliability latency SLAs, uptime, error budgets - Build evaluation and regression harnesses so accuracy and latency dont silently regress release to release Robustness real-time Voice AI - Build robust ASR for real-world environments noisy audio, telephony, far-field speech, reverberation, and overlapping speakers - Develop the Voice AI components real-time conversations depend on: VAD, endpointing, turn detection, barge-in, speech segmentation, and conversational state management - Turn recurring field failures (accent, domain, channel) into systematic model and pipeline improvements Technical leadership research presence - Technically guide engineers experiment design, tracking, code and research review, and career-shaping mentorship - Translate ASR research into enterprise-scale production systems and define the architecture that gets there - Set the bar for rigor: reproducible experiments, clean baselines, honest evaluation, and clear write-ups - Stay current with ASR research and bring the relevant frontier into Gnanis roadmap - Publish and represent Gnanis ASR work externally international conferences, talks, and community engagement Who you are EXPERIENCE WERE LOOKING FOR - 8 10 years in speech / ML research or engineering, with substantial contribution to production-grade ASR systems - A track record of designing, training, and deploying streaming and offline ASR at scale high accuracy, low latency, high reliability - Deep expertise in multilingual ASR, especially Indian languages code-switching, regional accents, dialects, pronunciation variation, and low-resource languages - Strong command of modern ASR architectures Conformer, E-Branchformer, RNN-T, CTC, Whisper / encoder-decoder, and Speech Language Models (SLMs) - Good publications at international conferences (e.g. Interspeech, ICASSP, NeurIPS, ICML, ACL) in speech / audio / ML - A proven record of translating ASR research into enterprise-scale production systems and mentoring high-performing engineering teams WHAT MAKES A STANDOUT CANDIDATE - Experience building ASR for hostile audio telephony, far-field, reverberation, and overlapping Res
More at Gnani Innovations