Source description
About the role
Research and development of next-generation Vision Language Models across pre-training, instruction tuning, reasoning, and agents.
Develop novel architectures and training methodologies for integrating visual understanding, language reasoning, and tool-use capabilities.
Research efficient multimodal learning techniques, including data-efficient training, long-context modeling, model modularity, and inference optimization.
Build and improve large-scale multimodal datasets, synthetic data generation pipelines, and evaluation benchmarks for VLM capabilities.
Investigate multimodal reasoning, agentic behavior, OCR, grounding, document understanding, chart understanding, and visual question answering capabilities.
Contribute to technical reports, research publications, and open-source software.
Represent MBZUAI at research conferences and industry events, showcasing advancements in multimodal foundation models and large-scale AI systems.
Mentor junior researchers and collaborate across teams to drive impactful research initiatives.
More at Institute of Foundation Models
Related open roles
AI Research Internship - WM
San Francisco Bay Area · Onsite
Research Scientist, Agentic Data & Benchmarking
San Francisco Bay Area · Onsite
Research Engineer - The Diffusion LLM Team
San Francisco Bay Area · Onsite
Research Scientist - Agents
United States · Onsite
Research Scientist - Speech/Audio Machine Learning
Paris · Onsite
Research Scientist - Reinforcement Learning
San Francisco Bay Area · Onsite
