Source description
About the role
Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization
Design and implement performance benchmarks and testing methodologies to evaluate application performance
Build tools to automate workload analysis, workload optimization, and other critical workflows
Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization
Support the team to develop appropriate kernels and systems for new model architectures and algorithms
Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC and deep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.
Perform all other duties as reasonably directed by the line manager that are commensurate with these functional objectives.
More at Institute of Foundation Models
Related open roles
Inference Optimization Intern – Performance Modeling
United States · Onsite
Eval360 - Error Analysis Engineer
San Francisco Bay Area · Onsite
Machine Learning Engineer – World Model
San Francisco Bay Area · Onsite
Research Engineer - Speech/Audio Machine Learning
Paris · Onsite
Machine Learning Infrastructure Engineer
United States · Onsite
Machine Learning Engineer – World Modeling
Dubai · Onsite
