Source description
About the role
Yotta Data Services offers a comprehensive suite of cloud, data center, and managed services designed to accelerate digital transformation for businesses of all sizes. With state-of-the-art infrastructure, cutting-edge AI capabilities, and a commitment to data sovereignty, we empower organizations to innovate securely and efficiently. Job Scope Bridge cutting-edge ML research with production-scale AI infrastructure. Total /Relevant Experience 7 Plus Years of experience. Key Responsibilities Track and analyze frontier research (distributed training, MoE, inference systems) Prototype and validate research ideas on large GPU clusters Publish internal/external whitepapers - Guide platform and product direction for future technology horizon Must-have skill Strong PyTorch and distributed ML experience Systems-aware research mindset Ability to move from paper code insigh Good-to-Have Skills Experience with multi-node GPU environments Hands-on experience with distributed training frameworks Working knowledge of the NVIDIA ecosystem (TensorRT, Triton, NeMo) Experience deploying and operating AI models at scale on Kubernetes clusters Familiarity with Slurm or other workload schedulers Qualifications Criteria PhD in ML/ Systems / HPC Behavioral Attributes: Art of skillful conversation Creativity & Problem Solving Learning on fly Business Acumen Dealing with ambiguity Building Trust Customer Focus Intellectual Horsepower (Functional Skills) Action Orientation & Accountability Prioritizing, Planning & organizing Listening, Sensing, Observing Building Collaborative Relationships
More at Yotta Infrastructure