Source description
About the role
As a highly skilled System Administrator with expertise in Linux systems, networking, and low-latency infrastructure environments, you will be responsible for deploying, managing, monitoring, and optimizing mission-critical infrastructure with a focus on performance, reliability, and high availability. Your role will involve working closely with infrastructure, networking, and engineering teams to maintain ultra-low latency environments, troubleshoot complex production issues, and continuously improve system performance. Key Responsibilities: - Deploy, configure, and maintain Linux servers across Ubuntu, Red Hat, and CentOS environments. - Perform Linux system tuning including CPU pinning, NUMA optimization, IRQ affinity, huge pages configuration, and kernel parameter tuning. - Manage OS-level scheduling policies, real-time priorities, and CPU isolation to improve system performance. - Handle system hardening, patch management, firmware upgrades, and OS lifecycle management. - Maintain infrastructure documentation, SOPs, and operational runbooks. - Manage and optimize bare-metal Linux infrastructure for high-performance environments. - Support infrastructure related to market data, order management, and execution systems. - Monitor and optimize latency, jitter, and packet loss across the infrastructure stack. - Work on high-speed network hardware and ultra-low latency networking environments. - Support exchange connectivity and colocation infrastructure operations. - Manage L2/L3 networking including VLANs, BGP, OSPF, multicast, and high-availability configurations. - Configure and manage Cisco Nexus or similar data center switches. - Coordinate with data center teams for rack management, cabling, power, and cooling activities. - Ensure redundancy, failover readiness, and proactive infrastructure capacity planning. - Build and maintain monitoring and alerting systems using Prometheus and Grafana. - Automate routine operational tasks using Bash and Python scripting. - Perform infrastructure health checks, deployment automation, log management, and troubleshooting. - Collaborate with engineering teams to resolve infrastructure bottlenecks and production incidents. Required Technical Skills: - Core Linux: Ubuntu, Red Hat, CentOS, Linux internals, Kernel tuning, CPU pinning, NUMA optimization, BIOS/hardware tuning. - Networking: TCP/IP, UDP, Multicast, VLANs, BGP, OSPF, L2/L3 networking. - Low-Latency & Infrastructure: High-speed networking (10/25/40/100GbE), Mellanox/NVIDIA NICs, Solarflare NICs, Cisco Nexus switches, DPDK/OpenOnload exposure, NVMe/high-performance storage systems. - Monitoring & Automation: Bash scripting, Python scripting, Prometheus, Grafana. If you meet the qualifications and possess the required technical skills, have experience in low-latency, HFT, trading, financial services, or high-performance computing environments, and demonstrate strong troubleshooting skills in mission-critical production infrastructure, this role could be an excellent fit for you. Your ability to work in fast-paced production environments, along with strong analytical and problem-solving capabilities, and good communication and coordination skills, will be valuable assets in this position. As a highly skilled System Administrator with expertise in Linux systems, networking, and low-latency infrastructure environments, you will be responsible for deploying, managing, monitoring, and optimizing mission-critical infrastructure with a focus on performance, reliability, and high availability. Your role will involve working closely with infrastructure, networking, and engineering teams to maintain ultra-low latency environments, troubleshoot complex production issues, and continuously improve system performance. Key Responsibilities: - Deploy, configure, and maintain Linux servers across Ubuntu, Red Hat, and CentOS environments. - Perform Linux system tuning including CPU pinning, NUMA optimization, IRQ affinity, huge pages configuration, and kernel parameter tuning. - Manage OS-level scheduling policies, real-time priorities, and CPU isolation to improve system performance. - Handle system hardening, patch management, firmware upgrades, and OS lifecycle management. - Maintain infrastructure documentation, SOPs, and operational runbooks. - Manage and optimize bare-metal Linux infrastructure for high-performance environments. - Support infrastructure related to market data, order management, and execution systems. - Monitor and optimize latency, jitter, and packet loss across the infrastructure stack. - Work on high-speed network hardware and ultra-low latency networking environments. - Support exchange connectivity and colocation infrastructure operations. - Manage L2/L3 networking including VLANs, BGP, OSPF, multicast, and high-availability configurations. - Configure and manage Cisco Nexus or similar data center switches. - Coordinate with data center teams for rack management, cabling, power, an
More at Whitefield Careers