Padmi

KGeN - Voice AI Researcher (India)

IndiaPosted 2 months ago
Computer ResearchJuniorFull Time; Regular
Apply at KGEN

Opens the source posting on shine.com

Source description

About the role

View original

Speech AI Research Engineer About the Role : We are building structured, high-quality voice datasets for frontier AI companies working on speech-to-text, speech-to-speech, and multimodal AI systems. We are looking for a Machine Learning Researcher with a focus on voice and speech AI - someone who can rigorously evaluate datasets across evolving speech models, identify performance gaps across Indic and global languages, and publish those findings as structured research for the broader AI community. This role sits at the intersection of benchmarking, linguistic diversity, and data strategy. If you are deeply curious about how models fail - especially across underrepresented languages and accents - this is built for you. What You'll Own : 1. Cross-Model Benchmarking & Evaluation : - Benchmark voice datasets across ASR and speech models (Whisper, Deepgram, Google STT, Azure Speech, and emerging open-source models). - Measure performance using WER, CER, MOS, robustness, latency, and error pattern analysis. - Design structured experiments to understand how dataset characteristics impact model accuracy. - Compare performance across multilingual, dialect-heavy, emotional, and noisy speech data. 2. Model Gap Analysis - Indic & Global Languages : - Systematically identify where speech models underperform across : Indic languages and dialects (Hindi, Tamil, Telugu, Bengali, Kannada, etc.), code-switching and transliteration, emotional and conversational speech, low-resource language scenarios, and background noise / real-world audio conditions. - Quantify model weaknesses through structured, reproducible analysis. - Map performance gaps to specific dataset requirements - you will help define what data models actually need next. 3. Dataset Quality & Supplier Scoring : - Build a standardized dataset quality scoring rubric with measurable Speech AI Research Engineer About the Role : We are building structured, high-quality voice datasets for frontier AI companies working on speech-to-text, speech-to-speech, and multimodal AI systems. We are looking for a Machine Learning Researcher with a focus on voice and speech AI - someone who can rigorously evaluate datasets across evolving speech models, identify performance gaps across Indic and global languages, and publish those findings as structured research for the broader AI community. This role sits at the intersection of benchmarking, linguistic diversity, and data strategy. If you are deeply curious about how models fail - especially across underrepresented languages and accents - this is built for you. What You'll Own : 1. Cross-Model Benchmarking & Evaluation : - Benchmark voice datasets across ASR and speech models (Whisper, Deepgram, Google STT, Azure Speech, and emerging open-source models). - Measure performance using WER, CER, MOS, robustness, latency, and error pattern analysis. - Design structured experiments to understand how dataset characteristics impact model accuracy. - Compare performance across multilingual, dialect-heavy, emotional, and noisy speech data. 2. Model Gap Analysis - Indic & Global Languages : - Systematically identify where speech models underperform across : Indic languages and dialects (Hindi, Tamil, Telugu, Bengali, Kannada, etc.), code-switching and transliteration, emotional and conversational speech, low-resource language scenarios, and background noise / real-world audio conditions. - Quantify model weaknesses through structured, reproducible analysis. - Map performance gaps to specific dataset requirements - you will help define what data models actually need next. 3. Dataset Quality & Supplier Scoring : - Build a standardized dataset quality scoring rubric with measurable

One address, no account. We’ll tell you when matching roles go live.

More at KGEN

Related open roles

View all roles