Padmi

HPC System engineer

IndiaPosted 3 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at Nexifyr Consulting Pvt Ltd

Opens the source posting on shine.com

Source description

About the role

View original

As an HPC Systems Engineer with 37 years of experience in HPC, you will be responsible for the following: - Diagnosing and resolving HPC issues related to applications, scheduler, and storage. - Analyzing job failures, performance bottlenecks, and system logs. - Managing and optimizing schedulers like SLURM and PBS. - Performing cluster health checks and proactive monitoring. - Installing and supporting HPC applications and user environments. - Troubleshooting networking issues, particularly InfiniBand and Ethernet. - Identifying recurring issues and implementing permanent fixes. - Collaborating with L3 for resolving deep technical issues. - Automating routine operational tasks using scripts. - Updating and improving standard operating procedures (SOPs), runbooks, and documentation. Required Skills: - Deep understanding of HPC architecture - Strong Linux administration skills - Experience with SLURM and PBS - Knowledge of parallel computing (MPI, OpenMP) - Familiarity with storage systems such as Lustre, NFS, and GPFS - Networking knowledge, InfiniBand preferred - Proficiency in scripting languages like Python and Bash - Knowledge of AWS ParallelCluster and AWS PCS will be an advantage (Note: No additional details about the company were mentioned in the job description) As an HPC Systems Engineer with 37 years of experience in HPC, you will be responsible for the following: - Diagnosing and resolving HPC issues related to applications, scheduler, and storage. - Analyzing job failures, performance bottlenecks, and system logs. - Managing and optimizing schedulers like SLURM and PBS. - Performing cluster health checks and proactive monitoring. - Installing and supporting HPC applications and user environments. - Troubleshooting networking issues, particularly InfiniBand and Ethernet. - Identifying recurring issues and implementing permanent fixes. - Collaborating with L3 for resolving deep technical issues. - Automating routine operational tasks using scripts. - Updating and improving standard operating procedures (SOPs), runbooks, and documentation. Required Skills: - Deep understanding of HPC architecture - Strong Linux administration skills - Experience with SLURM and PBS - Knowledge of parallel computing (MPI, OpenMP) - Familiarity with storage systems such as Lustre, NFS, and GPFS - Networking knowledge, InfiniBand preferred - Proficiency in scripting languages like Python and Bash - Knowledge of AWS ParallelCluster and AWS PCS will be an advantage (Note: No additional details about the company were mentioned in the job description)

One address, no account. We’ll tell you when matching roles go live.

More at Nexifyr Consulting Pvt Ltd

Related open roles

View all roles