Source description
About the role
Understand, analyze, profile, optimize, and provide guidance to the team on deep learning workloads on state-of-the-art hardware and software platforms to improve their efficiency with different levels of optimization
Design and implement performance benchmarks and testing methodologies to evaluate application performance
Build tools to automate workload analysis, workload optimization, and other critical workflows
Triage system issues and identify bottleneck and inefficiencies by analyzing the sources of issues and the impact on hardware, network and propose solutions to enhance GPU utilization
Support the team to develop appropriate kernels and systems for new model architectures and algorithms
Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
Review code developed by other developers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
Represent MBZUAI at industry conferences and events, showcasing the institution’s cutting-edge HPC and deep learning capabilities and establishing MBZUAI as a global leader in AI research and innovation.
Perform all other duties as reasonably directed by the line manager that are commensurate with these functional objectives.
More at Institute of Foundation Models
Related open roles
Eval360 - Error Analysis Engineer
San Francisco Bay Area · Onsite
Senior Distributed Systems Engineer
San Francisco Bay Area · Onsite
Senior MLOps Engineer
Dubai · Onsite
High Performance Computing Software Engineer - Supercomputing
Dubai · Onsite
Machine Learning Engineer – World Modeling
Dubai · Onsite
Data Engineer
Dubai · Onsite
