Padmi

IT Systems Administrator 4 (HPC Cluster Administrator) (Bengaluru)

BangalorePosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Kairahire Solutions

Opens the source posting on shine.com

Source description

About the role

View original

Location: Bangalore - White field Work Mode: Onsite(all 5 days) Exp: 6 to 8 years of relevant exp in Linux Admin, HPC Cluster Top 3 Skills: 6 to 8 years of relevant exp in Linux Admin, HPC Cluster, o Proficiency in scripting languages like Python, Bash, or Ansible. o Experience in debugging and maintaining Slurm based HPC Cluster, o Solid analytical and troubleshooting skills. Good communication High Performance Computing (HPC) Cluster Administrator At AMD, we push the boundaries of what is possible. Key Responsibilities - System Administration: Install, configure, and maintain Linux operating systems (RHEL, CentOS, SLES, Ubuntu, etc.) on all cluster nodes. Cluster Management: Deploy, manage, and troubleshoot HPC cluster management tools and environments, including physical hardware. Performance Optimization: Monitor system health and utilization, tuning the environment to maximize performance and ensure optimal resource allocation. Job Scheduling and Resource Management: Administer workload management and job scheduling systems (e.g., Slurm). Storage and Networking: Manage large-scale parallel file systems (Beegfs, GPFS, etc.) and high-speed networking technologies (InfiniBand, high-speed Ethernet, RoCE). Software Support: Install, upgrade, compile, and maintain scientific software, applications, and libraries (e.g., MPI, compilers, debuggers) required by users. Automation and Scripting: Develop and automate administrative tasks and streamline operations via Ansible. User Support and Collaboration: Provide technical support to users, troubleshoot issues, and create comprehensive documentation and user guides. Security and Compliance: Ensure the system security posture is maintained, adhering to security standards and regulatory requirements. Containers and Micro-services: Enable a containerized environment in the HPC cluster and convert critical services to micro-services and containerized scientific applications like Gromacs, Openfoam, cp2k, etc. Compensation: 1,000,000.00 - 1,700,000.00 per year Benefits - Flexible schedule - Health insurance - Leave encashment - Paid sick time - Provident Fund Work Location: In person .

One address, no account. We’ll tell you when matching roles go live.

More at Kairahire Solutions

Related open roles

View all roles