Source description
About the role
Technical Excellence
- [REQUIRED] 5+ years of experience in deep learning research with focus on large-scale model training
- [REQUIRED] Demonstrated experience training models at scale (100M+ parameters) from scratch
- [REQUIRED] Deep understanding of transformer architectures, attention mechanisms, and modern training techniques (mixed precision, distributed training, gradient accumulation)
- Production experience with ML systems at scale (PyTorch, distributed training, model serving)
- Experience with language model pre-training, including tokenization strategies, training objectives (CLM, MLM), and scaling laws
- Understanding of fine-tuning techniques including instruction tuning, RLHF, and preference optimization (DPO, PPO)
- Knowledge of efficient training techniques: LoRA, QLoRA, flash attention, gradient checkpointing
- Experience with model evaluation, benchmarking, and safety considerations
Systems Experience
- Background in vision-language models or multimodal architectures strongly preferred
- Experience building and optimizing data pipelines for large-scale training
- Systems-level thinking about training efficiency, hardware utilization, and cost optimization
- Familiarity with MLOps for model versioning, experiment tracking, and reproducibility
Preferred Qualifications
- Advanced degree (M.S./Ph.D.) in Computer Science, Electrical Engineering, or related field with focus on deep learning
- First-author publications at top ML venues (NeurIPS, ICML, ICLR, ACL, CVPR) on language models, multimodal learning, or efficient training
- Experience training or contributing to open-source foundation models
- Background in domain-specific model development (code, science, medical, etc.)
- Experience with video understanding models or temporal reasoning in transformers
- Contributions to major ML frameworks or training libraries
- Track record of transitioning research prototypes to production systems
Location & Compensation
- San Francisco Bay Area (on-site)
- Competitive salary and significant equity package
- Full benefits including health, dental, vision, and 401k +6% match
- Access to dedicated GPU compute resources for research and experimentation
Compensation
The base pay range for this role is $180,000 – $350,000 per year.
Ready to apply?
Powered by
First name *
Last name *
Email *
Resume *
Click to upload or drag and drop here
LinkedIn URL
Apply
Req ID: R2D5
More at Ironsite AI
Related open roles
IT Engineer
San Francisco Bay Area · Onsite
Electrical Engineer
San Francisco Bay Area · Onsite
$150k–$300k/yr
Member of Technical Staff - Platform
San Francisco Bay Area · Onsite
$150k–$300k/yr
Applied ML Researcher
San Francisco Bay Area · Onsite
$180k–$350k/yr
Mechanical Engineer
San Francisco Bay Area · Onsite
$150k–$300k/yr
Member of Technical Staff - Infrastructure
San Francisco Bay Area · Onsite
$175k–$325k/yr
