Source description
About the role
Role Overview: NVIDIA is seeking a Senior AI/HPC Engineer to join its infrastructure Specialist team. You will have the opportunity to work on dynamic, customer-focused projects that involve deploying, managing, and maintaining AI/HPC infrastructure in Linux-based environments. This role requires excellent interpersonal skills as you will be interacting with customers, partners, and internal teams to analyze, define, and implement large-scale AI/HPC projects. Key Responsibilities: - Deploy, manage, and maintain AI/HPC infrastructure in Linux-based environments for new and existing customers. - Act as the domain expert with customers during planning calls through implementation. - Provide handover-related documentation and perform knowledge transfers to support customers as they begin deploying sophisticated systems. - Offer feedback to internal teams by reporting bugs, documenting workarounds, and suggesting improvements. Qualifications Required: - BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. - 5+ years of experience providing in-depth support and deployment services for hardware and software products. - Knowledge and experience with Linux System Administration, including process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, and network routing/advanced networking. - Proficiency in scripting and cluster management technologies. - Strong interpersonal skills with the ability to resolve customer-blocking issues effectively. - Excellent verbal and written English skills. - Strong organizational skills with the ability to prioritize and multitask with limited supervision. - Industry-standard Linux certifications. - Experience with Schedulers such as SLURM, LSF, UGE, etc. - Hands-on experience with MPI, NCCL principles, high-speed networks, automation tools, and Kubernetes will be advantageous. NVIDIA, known for its innovation and talented workforce, offers an equal opportunity work environment that values diversity and inclusivity. If you are a creative and autonomous manager looking to join a leading SW design team, NVIDIA could be the right place for you to make an impact. Role Overview: NVIDIA is seeking a Senior AI/HPC Engineer to join its infrastructure Specialist team. You will have the opportunity to work on dynamic, customer-focused projects that involve deploying, managing, and maintaining AI/HPC infrastructure in Linux-based environments. This role requires excellent interpersonal skills as you will be interacting with customers, partners, and internal teams to analyze, define, and implement large-scale AI/HPC projects. Key Responsibilities: - Deploy, manage, and maintain AI/HPC infrastructure in Linux-based environments for new and existing customers. - Act as the domain expert with customers during planning calls through implementation. - Provide handover-related documentation and perform knowledge transfers to support customers as they begin deploying sophisticated systems. - Offer feedback to internal teams by reporting bugs, documenting workarounds, and suggesting improvements. Qualifications Required: - BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. - 5+ years of experience providing in-depth support and deployment services for hardware and software products. - Knowledge and experience with Linux System Administration, including process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, and network routing/advanced networking. - Proficiency in scripting and cluster management technologies. - Strong interpersonal skills with the ability to resolve customer-blocking issues effectively. - Excellent verbal and written English skills. - Strong organizational skills with the ability to prioritize and multitask with limited supervision. - Industry-standard Linux certifications. - Experience with Schedulers such as SLURM, LSF, UGE, etc. - Hands-on experience with MPI, NCCL principles, high-speed networks, automation tools, and Kubernetes will be advantageous. NVIDIA, known for its innovation and talented workforce, offers an equal opportunity work environment that values diversity and inclusivity. If you are a creative and autonomous manager looking to join a leading SW design team, NVIDIA could be the right place for you to make an impact.
More at NVIDIA
Related open roles
Senior Site Reliability Engineer, Senior Site Reliability Engineer
Mumbai
HPC Infra Engineer (Hyderabad)
Hyderabad
Compute Cluster SRE Engineer, GPU - HPC (Bengaluru)
Bangalore
Senior Platform and EngOps Engineer - Cluster Operations
Bangalore
Senior DevOps Engineer - E-commerce
Mumbai
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
India