Padmi
Institute of Foundation Models logo
Institute of Foundation Models

foundation models · large language models

Research Scientist - Vision Language Model

United States · Onsite$150k–$450k/yrPosted 2 months ago
AI researchUnspecifiedFull Time
Apply at Institute of Foundation Models

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Research and development of next-generation Vision Language Models across pre-training, instruction tuning, reasoning, and agents.

Develop novel architectures and training methodologies for integrating visual understanding, language reasoning, and tool-use capabilities.

Research efficient multimodal learning techniques, including data-efficient training, long-context modeling, model modularity, and inference optimization.

Build and improve large-scale multimodal datasets, synthetic data generation pipelines, and evaluation benchmarks for VLM capabilities.

Investigate multimodal reasoning, agentic behavior, OCR, grounding, document understanding, chart understanding, and visual question answering capabilities.

Contribute to technical reports, research publications, and open-source software.

Represent MBZUAI at research conferences and industry events, showcasing advancements in multimodal foundation models and large-scale AI systems.

Mentor junior researchers and collaborate across teams to drive impactful research initiatives.

More at Institute of Foundation Models

Related open roles

View all roles