Source description
About the role
Job Summary We are seeking an experienced Senior Big Data Platform Administrator to manage and support our enterprise Big Data infrastructure. The ideal candidate will have strong expertise in the Hadoop ecosystem, Linux administration, distributed systems, and platform monitoring tools. This role involves ensuring platform performance, availability, scalability, and security across production and non-production environments. Key Responsibilities Administer and support Hadoop ecosystem components (HDFS, YARN) across Production, Test, and Development environments Manage and maintain Apache Spark and Airflow infrastructure Monitor cluster health, troubleshoot performance issues, and optimize system performance Implement and maintain High Availability (HA), Disaster Recovery (DR), backup, and restore strategies Manage logging, monitoring, and alerting systems using Prometheus, Grafana, and ELK Stack Perform Linux (Ubuntu) system administration, including patching, upgrades, and security hardening Develop and maintain automation scripts using Shell and/or Python Handle Kerberos authentication setup and maintenance Support incident management, root cause analysis (RCA), and ensure SLA adherence Manage capacity planning and system scalability Follow CI/CD and DevOps best practices for platform improvements and deployments Job Summary We are seeking an experienced Senior Big Data Platform Administrator to manage and support our enterprise Big Data infrastructure. The ideal candidate will have strong expertise in the Hadoop ecosystem, Linux administration, distributed systems, and platform monitoring tools. This role involves ensuring platform performance, availability, scalability, and security across production and non-production environments. Key Responsibilities Administer and support Hadoop ecosystem components (HDFS, YARN) across Production, Test, and Development environments Manage and maintain Apache Spark and Airflow infrastructure Monitor cluster health, troubleshoot performance issues, and optimize system performance Implement and maintain High Availability (HA), Disaster Recovery (DR), backup, and restore strategies Manage logging, monitoring, and alerting systems using Prometheus, Grafana, and ELK Stack Perform Linux (Ubuntu) system administration, including patching, upgrades, and security hardening Develop and maintain automation scripts using Shell and/or Python Handle Kerberos authentication setup and maintenance Support incident management, root cause analysis (RCA), and ensure SLA adherence Manage capacity planning and system scalability Follow CI/CD and DevOps best practices for platform improvements and deployments
More at CIEL HR