Source description
About the role
As an HPC Systems Engineer with 37 years of experience in HPC, you will be responsible for the following: - Diagnosing and resolving HPC issues related to applications, scheduler, and storage. - Analyzing job failures, performance bottlenecks, and system logs. - Managing and optimizing schedulers like SLURM and PBS. - Performing cluster health checks and proactive monitoring. - Installing and supporting HPC applications and user environments. - Troubleshooting networking issues, particularly InfiniBand and Ethernet. - Identifying recurring issues and implementing permanent fixes. - Collaborating with L3 for resolving deep technical issues. - Automating routine operational tasks using scripts. - Updating and improving standard operating procedures (SOPs), runbooks, and documentation. Required Skills: - Deep understanding of HPC architecture - Strong Linux administration skills - Experience with SLURM and PBS - Knowledge of parallel computing (MPI, OpenMP) - Familiarity with storage systems such as Lustre, NFS, and GPFS - Networking knowledge, InfiniBand preferred - Proficiency in scripting languages like Python and Bash - Knowledge of AWS ParallelCluster and AWS PCS will be an advantage (Note: No additional details about the company were mentioned in the job description) As an HPC Systems Engineer with 37 years of experience in HPC, you will be responsible for the following: - Diagnosing and resolving HPC issues related to applications, scheduler, and storage. - Analyzing job failures, performance bottlenecks, and system logs. - Managing and optimizing schedulers like SLURM and PBS. - Performing cluster health checks and proactive monitoring. - Installing and supporting HPC applications and user environments. - Troubleshooting networking issues, particularly InfiniBand and Ethernet. - Identifying recurring issues and implementing permanent fixes. - Collaborating with L3 for resolving deep technical issues. - Automating routine operational tasks using scripts. - Updating and improving standard operating procedures (SOPs), runbooks, and documentation. Required Skills: - Deep understanding of HPC architecture - Strong Linux administration skills - Experience with SLURM and PBS - Knowledge of parallel computing (MPI, OpenMP) - Familiarity with storage systems such as Lustre, NFS, and GPFS - Networking knowledge, InfiniBand preferred - Proficiency in scripting languages like Python and Bash - Knowledge of AWS ParallelCluster and AWS PCS will be an advantage (Note: No additional details about the company were mentioned in the job description)
More at Nexifyr Consulting Pvt Ltd