Padmi

Senior ML Engineer Multimodal Video Generation & LLMs

IndiaPosted 1 month ago
Computer ResearchSeniorFull Time; Regular
Apply at NGFLIX.COM

Opens the source posting on shine.com

Source description

About the role

View original

About NGFLIX.COM NGFLIX.COM is an AI-powered content marketplace, OTT platform, and creative studio focused on transforming how video content is created, shared, and experienced. The company integrates advanced AI technologies to help creators develop high-quality video experiences with greater speed and creative freedom. NGFLIX.COM offers tools for both watching major releases and generating original cinematic content. The platform is designed for creators, technologists, and viewers who want to push the boundaries of storytelling through AI-driven innovation. Role Description This is a full-time remote role for a Senior ML Engineer Multimodal Video Generation & LLMs. The role involves designing, training, and optimizing machine learning models for video generation, multimodal understanding, and large language model integration. Day-to-day responsibilities include building end-to-end ML pipelines, experimenting with new architectures for generative and multimodal models, and improving performance, latency, and scalability in production systems. The Senior ML Engineer collaborates closely with product, engineering, and content teams to translate creative and business requirements into robust technical solutions and conducts rigorous experimentation, evaluation, and model monitoring. The ideal candidate has deep experience working with open-source frameworks (Hugging Face, Diffusers, PyTorch) and integrating large language models (LLMs) into multimodal pipelines for video generation. Required Qualifications Bachelors/masters in computer science, AI, or related field with 5+ years of experience. Strong programming skills in Python, CUDA, and C++/Rust. Expertise in Neural Networks and Pattern Recognition, particularly in deep learning architectures for vision, sequence modeling, and multimodal tasks. Build and optimize text-to-video generation models from scratch using multimodal inputs (text, audio, image, video). Exposure to open-source generative AI projects like Wan-AI, HunyuanVideo, Genmo Mochi, SulphurAI, LongCat-Video. Architect fusion mechanisms for LLM-driven semantic alignment in video generation. Implement evaluation metrics for video realism, coherence, and text alignment. Perks & Benefits Competitive Compensation Work Flexibility Equity Options .

One address, no account. We’ll tell you when matching roles go live.