Source description
About the role
Job Title : Senior Linux Systems Engineer High Performance Computing (HPC) Location : Chennai - Hybrid Job Type : Permanent Role Experience Required : 6+ Years in IT (Minimum 4+ Years in Active Infrastructure Automation or Systems Development) Notice Period : Immediate to 15 Days Max Education Required : Bachelor's Degree in IT/Engineering Job Summary : We are seeking a highly skilled and motivated Linux Systems Engineer to join our global HPC Supercomputing platform team. In this role, you will focus on the design, optimization, and scaling of core Computer-Aided Engineering (CAE) simulation, graphical rendering, and analytics capabilities used by our global product development groups. The ideal candidate is a seasoned Level 3 systems engineer who bridges the gap between low-level Linux kernel/network structures and high-concurrency scientific computing. You will be responsible for profiling and tuning heavy distributed simulations, integrating scalable CAE applications across thousands of compute nodes, building custom command-line interface (CLI) tooling and APIs for engineer consumption, and identifying systemic architecture bottlenecks through advanced telemetry layers. Key Responsibilities : - HPC Workload Optimization & Tuning : Install, profile, benchmark, and tune core CAE application workflows and highly distributed workloads to maximize compute utilization across the HPC platform. - Hardware Evaluation : Test and evaluate demanding scientific workloads across the latest CPU and GPU microarchitectures to drive hardware choices and ensure cost-efficient simulation delivery. - Tooling & API Development : Design and program custom CLI tools, utilities, and inner APIs (using Python, Go, or Bash) that internal engineering teams consume to streamline access to supercomputing infrastructure. - Parallel Architecture Integration : Manage, monitor, and scale distributed communication layers and Message Passing Interface standards (IntelMPI, OpenMPI, etc.) across clustered network paths. - Cluster Management & Batch Scheduling : Configure and manage automated environment state lines and job queues using batch schedulers (such as PBS Pro or Slurm). - Infrastructure as Code (IaC) & Containerization : Provision, automate, and configure cluster nodes via configuration management utilities (Ansible, Puppet, or Chef) and deploy secure compute containers using Docker, Apptainer (Singularity), or Kubernetes. - Deep Telemetry & Observability : Architect platform telemetry, cluster dashboards, and resource tracking pipelines using metric collection and visualization engines like Prometheus and Grafana to identify system limits. - Technical Troubleshooting : Investigate, diagnose, and resolve complex structural failures across Linux kernels, high-speed fabrics, clustered storage volumes, and scientific runtime configurations. Skills & Qualifications : Required Technical DNA : - Operating Systems & Tuning : Strong, production-level engineering and command of Linux operating systems, ideally inside an HPC or high-concurrency server cluster environment. - Automation & Coding : Advanced coding and systems scripting proficiency in Python, Go, or Bash. - Parallel Workloads : Solid familiarity with distributed workloads, network topologies, and Message Passing Interface (MPI) frameworks. - HPC Applications : Hands-on operational experience supporting, profiling, or writing automation for CAE or parallel scientific analysis tools. - Tenure : 6+ years of total IT experience, with at least 4 years spent in active infrastructure automation, product delivery, or system software development. Preferred Qualifications (Nice to Have) : - Queue Managers : Hands-on experience configuring batch engines like PBS Pro or Slurm. - Configuration Management : Production experience automating bare-metal or cloud platforms using Ansible. - Technical utility deploying modern telemetry systems (Prometheus, Grafana) and container nodes (Apptainer, Docker). .
More at The Judge Group