Source description
About the role
At Humyn Labs, the best AI is built on the best human judgment. We operate a global network of 1M+ verified experts delivering high-quality, multimodal training datasets across various domains. Our data is thoroughly evaluated, defended, and made production-ready to ensure the trustworthiness of AI. We focus on egocentric video understanding, embodied AI, robotics perception, and voice-driven interaction. We prioritize data quality, scalability, and fast-paced innovation. Key Responsibilities: - Benchmark voice datasets across ASR and speech models like Whisper, Deepgram, Google STT, Azure Speech, and emerging open-source models. - Measure performance using WER, CER, MOS, robustness, latency, and error pattern analysis. - Design experiments to understand how dataset characteristics impact model accuracy. - Compare performance across multilingual, dialect-heavy, emotional, and noisy speech data. - Systematically identify underperformance areas in speech models across Indic languages and dialects, emotional speech, low-resource scenarios, and background noise. - Quantify model weaknesses through reproducible analysis and map performance gaps to specific dataset requirements. - Build a standardized dataset quality scoring rubric with measurable criteria and tag data suppliers based on quality signals. - Publish benchmarking findings as blog posts and LinkedIn articles accessible to technical and non-technical audiences. - Contribute to internal evaluation reports tracking performance shifts with new models. - Stay updated on evolving speech model architectures and share insights with research teams and clients. Qualifications Required: - 13 years of experience in speech AI, audio ML, NLP, or applied AI research. - Hands-on experience with ASR/TTS systems and understanding of model behavior. - Exposure to running experiments, evaluating models, and designing evaluation frameworks. - Strong Python skills and comfort with ML experimentation workflows. - Genuine interest in linguistic diversity, especially in Indic languages, and model performance across them. - Strong written communication skills to convert research into clear, publishable content. Additional Details: Ideal Mindset: - Curious about model failure modes, not just capabilities. - Analytical and detail-oriented with a bias for reproducibility. - Comfortable reading research papers and independently testing new APIs. - Excited to share work publicly through blogs, LinkedIn, and open datasets. (Note: Omitted additional details section as it was not present in the JD) At Humyn Labs, the best AI is built on the best human judgment. We operate a global network of 1M+ verified experts delivering high-quality, multimodal training datasets across various domains. Our data is thoroughly evaluated, defended, and made production-ready to ensure the trustworthiness of AI. We focus on egocentric video understanding, embodied AI, robotics perception, and voice-driven interaction. We prioritize data quality, scalability, and fast-paced innovation. Key Responsibilities: - Benchmark voice datasets across ASR and speech models like Whisper, Deepgram, Google STT, Azure Speech, and emerging open-source models. - Measure performance using WER, CER, MOS, robustness, latency, and error pattern analysis. - Design experiments to understand how dataset characteristics impact model accuracy. - Compare performance across multilingual, dialect-heavy, emotional, and noisy speech data. - Systematically identify underperformance areas in speech models across Indic languages and dialects, emotional speech, low-resource scenarios, and background noise. - Quantify model weaknesses through reproducible analysis and map performance gaps to specific dataset requirements. - Build a standardized dataset quality scoring rubric with measurable criteria and tag data suppliers based on quality signals. - Publish benchmarking findings as blog posts and LinkedIn articles accessible to technical and non-technical audiences. - Contribute to internal evaluation reports tracking performance shifts with new models. - Stay updated on evolving speech model architectures and share insights with research teams and clients. Qualifications Required: - 13 years of experience in speech AI, audio ML, NLP, or applied AI research. - Hands-on experience with ASR/TTS systems and understanding of model behavior. - Exposure to running experiments, evaluating models, and designing evaluation frameworks. - Strong Python skills and comfort with ML experimentation workflows. - Genuine interest in linguistic diversity, especially in Indic languages, and model performance across them. - Strong written communication skills to convert research into clear, publishable content. Additional Details: Ideal Mindset: - Curious about model failure modes, not just capabilities. - Analytical and detail-oriented with a bias for reproducibility. - Comfortable reading research papers and indepen