Padmi

Scribie- Senior Applied ML Engineer Audio and Foundation Models

BangalorePosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Scribie

Opens the source posting on shine.com

Source description

About the role

View original

Scribie is an AI-powered, Human Verified audio and video transcription service, trusted globally since 2008. We specialize in delivering accurate and reliable transcription solutions by blending advanced AI technology with human expertise. Headquartered in the US, we operate with a hybrid model in our Bangalore office, combining the flexibility of remote work with the collaboration of in-person engagement. This approach offers our team both autonomy and growth opportunities in a dynamic and supportive environment.Were building production-grade audio foundation models for high-stakes legal and enterprise transcription real customer data, messy audio, real consequences.This is not a paper-only research role.Youll own the full ML lifecycle:Fine-tuning large audio / multimodal models using SFT, LoRA, and RL-based preference optimization (DPO / PPO / ORPO)Beating strong baselines like Whisper-large, GPT-4o, Gemini, Claude on domain-specific dataDesigning WER, diarization, and alignment-driven evaluation stacksTaking models from research notebooks production inference servicesRunning daily experiments that directly impact quality, cost, and customer satisfactionYoull be our first ML hire, with real ownership over the audio ML roadmap not a side project, not a support role. Bangalore (primarily onsite) Compensation - 25L 30LThis role is a great fit if you:Have shipped fine-tuned ASR / LLM / multimodal models into productionAre comfortable running large training jobs and debugging failuresCare about real-world impact, not just benchmarksNot a fit if youre looking for an academic or paper-only research role.Apply here Or DM me with your LinkedIn/GitHub and a short note on the coolest audio or LLM system youve shipped.If turning messy real-world audio into models that make humans 510 more efficient excites you lets talk. Scribie is an AI-powered, Human Verified audio and video transcription service, trusted globally since 2008. We specialize in delivering accurate and reliable transcription solutions by blending advanced AI technology with human expertise. Headquartered in the US, we operate with a hybrid model in our Bangalore office, combining the flexibility of remote work with the collaboration of in-person engagement. This approach offers our team both autonomy and growth opportunities in a dynamic and supportive environment.Were building production-grade audio foundation models for high-stakes legal and enterprise transcription real customer data, messy audio, real consequences.This is not a paper-only research role.Youll own the full ML lifecycle:Fine-tuning large audio / multimodal models using SFT, LoRA, and RL-based preference optimization (DPO / PPO / ORPO)Beating strong baselines like Whisper-large, GPT-4o, Gemini, Claude on domain-specific dataDesigning WER, diarization, and alignment-driven evaluation stacksTaking models from research notebooks production inference servicesRunning daily experiments that directly impact quality, cost, and customer satisfactionYoull be our first ML hire, with real ownership over the audio ML roadmap not a side project, not a support role. Bangalore (primarily onsite) Compensation - 25L 30LThis role is a great fit if you:Have shipped fine-tuned ASR / LLM / multimodal models into productionAre comfortable running large training jobs and debugging failuresCare about real-world impact, not just benchmarksNot a fit if youre looking for an academic or paper-only research role.Apply here Or DM me with your LinkedIn/GitHub and a short note on the coolest audio or LLM system youve shipped.If turning messy real-world audio into models that make humans 510 more efficient excites you lets talk.

One address, no account. We’ll tell you when matching roles go live.

More at Scribie

Related open roles

View all roles