Source description
About the role
As a High-Performance Computing Network Engineer at our company, you will be responsible for ensuring the overall health and maintenance of storage technologies in our managed services customer's environments. You will be a valuable member of the Managed Services Infrastructure Practice, handling Tier 3 incident management, service request management, and change management infrastructure support for all Managed Services customers. Key Responsibilities: - Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities - Plan and perform maintenance activities - Assess customer environments for performance and design issues and propose resolutions - Work across technical teams to troubleshoot complex infrastructure issues - Create and maintain detailed documentation - Serve as a subject matter expert and escalation point for storage technologies - Collaborate with vendors to resolve storage issues - Communicate transparently with customers and internal team - Participate in on-call rotation - Complete training and certification assignments to enhance skills and knowledge Qualifications Required: - Bachelor's degree or equivalent in Information Systems or related field - 5+ years of expert level experience managing Network infrastructure in high-performance computing environments - Experience configuring, maintaining, and troubleshooting Nvidia/Mellanox (Cumulus OS) switches - Strong knowledge of Kubernetes and its networking components (CNI, Service Mesh, etc.) - Understanding of VPNs, Load Balancers, VPCs, and hybrid cloud networking - Experience with both ethernet and InfiniBand networking - Familiarity with high-performance computing (HPC) schedulers (e.g., SLURM, PBS, Torque) and their interaction with data storage systems - Experience with network containerization (Docker, Singularity) in an HPC context for data processing and application deployment - Solid working knowledge of Linux and Python scripting - Previous experience with network automation tools such as Ansible, Puppet, or Chef - Experience with machine learning or data science workflows in HPC environments - Managed Services or consulting experience - Strong background in customer service - High level problem-solving and communication skills - Strong oral and written communications skills - Related network certifications are a bonus Please note the compensation range indicated in the job posting reflects the On-Target Earnings (OTE) for this role, which includes a base salary and any applicable target bonus amount. This OTE range may vary based on your relevant experience, qualifications, and geographic location. As a High-Performance Computing Network Engineer at our company, you will be responsible for ensuring the overall health and maintenance of storage technologies in our managed services customer's environments. You will be a valuable member of the Managed Services Infrastructure Practice, handling Tier 3 incident management, service request management, and change management infrastructure support for all Managed Services customers. Key Responsibilities: - Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities - Plan and perform maintenance activities - Assess customer environments for performance and design issues and propose resolutions - Work across technical teams to troubleshoot complex infrastructure issues - Create and maintain detailed documentation - Serve as a subject matter expert and escalation point for storage technologies - Collaborate with vendors to resolve storage issues - Communicate transparently with customers and internal team - Participate in on-call rotation - Complete training and certification assignments to enhance skills and knowledge Qualifications Required: - Bachelor's degree or equivalent in Information Systems or related field - 5+ years of expert level experience managing Network infrastructure in high-performance computing environments - Experience configuring, maintaining, and troubleshooting Nvidia/Mellanox (Cumulus OS) switches - Strong knowledge of Kubernetes and its networking components (CNI, Service Mesh, etc.) - Understanding of VPNs, Load Balancers, VPCs, and hybrid cloud networking - Experience with both ethernet and InfiniBand networking - Familiarity with high-performance computing (HPC) schedulers (e.g., SLURM, PBS, Torque) and their interaction with data storage systems - Experience with network containerization (Docker, Singularity) in an HPC context for data processing and application deployment - Solid working knowledge of Linux and Python scripting - Previous experience with network automation tools such as Ansible, Puppet, or Chef - Experience with machine learning or data science workflows in HPC environments - Managed Services or consulting experience - Strong background in customer service - High level problem-solving a
More at AHEAD
