Source description
About the role
Role Summary NexTurn Inc. is seeking a highly skilled Red Hat Platform Engineer with deep hands-on expertise in OpenShift Container Platform (OCP) and OpenStack Platform (OSP). The ideal candidate will design, deploy, operate, and troubleshoot enterprise-grade containerized and virtualized infrastructure at scale, ensuring high availability, security, and performance across hybrid cloud environments. This role is critical in supporting client engagements that demand robust platform engineering, proactive monitoring, and rapid incident resolution across Red Hat technology stacks. Key Responsibilities OpenShift Container Platform (OCP) Deploy, configure, and manage OpenShift clusters (3.x / 4.x) in on-premises and hybrid cloud environments Perform cluster lifecycle management: upgrades, patching, scaling, node provisioning, and decommissioning Design and manage OpenShift Operators, OperatorHub integrations, and custom resource definitions (CRDs) Configure and maintain OpenShift networking SDN/OVN-Kubernetes, Ingress controllers, Routes, and Network Policies Administer RBAC, SCC (Security Context Constraints), OAuth integration, and multi-tenancy policies Integrate OCP with CI/CD pipelines (Tekton, Jenkins, GitLab CI) and internal developer platforms Manage persistent storage backends: Ceph/OCS, NFS, ODF, and dynamic provisioning via StorageClasses Monitor cluster health using Prometheus, Alertmanager, Grafana, and OpenShift built-in observability tools OpenStack Platform (OSP) Deploy and manage Red Hat OpenStack Platform using director/TripleO or Ansible-based deployment methods Administer core OSP services: Nova, Neutron, Cinder, Glance, Keystone, Heat, Swift, and Horizon Design and manage virtual networking VLANs, VXLAN overlays, provider networks, floating IPs, and security groups Configure and optimize compute flavors, hypervisor tuning (KVM/QEMU), CPU pinning, and NUMA topology Manage OSP storage tiers: Ceph integration, volume types, snapshot policies, and backup configurations Perform OSP upgrades, minor updates, and hotfix applications across overcloud and undercloud environments Implement identity federation, LDAP/AD integration, and quota/resource governance across projects Platform Troubleshooting & Incident Management Lead L2/L3 troubleshooting for OCP and OSP platform incidents including node failures, network disruptions, etcd issues, and scheduler anomalies Diagnose and resolve pod scheduling failures, CrashLoopBackOff, OOMKilled events, and image pull errors Troubleshoot Neutron networking faults DHCP failures, routing loops, SDN packet loss, and MTU mismatches Analyze and remediate storage I/O bottlenecks, PVC binding failures, and Ceph cluster degradation Perform RCA (Root Cause Analysis) and author post-incident reports with remediation action plans Engage Red Hat TAM/support channels with SOS reports, must-gather bundles, and case management Participate in on-call rotation for critical platform incidents; SLA adherence and escalation management Automation, DevOps & Tooling Develop and maintain Ansible playbooks and Roles for platform automation, configuration drift management, and remediation Implement GitOps workflows using ArgoCD/Flux for declarative cluster configuration management Write and maintain Bash/Python scripts for operational tooling, health checks, and reporting Manage infrastructure-as-code using Terraform or Ansible for cloud and on-prem resource provisioning Integrate with ServiceNow, Jira, or equivalent ITSM platforms for change, incident, and problem management Required Technical Skills Red Hat OCP 4.x Red Hat OSP 16/17 Kubernetes TripleO / Director Ansible / RHEL OVN-Kubernetes Ceph / ODF / OCS Prometheus / Grafana Tekton / Jenkins ArgoCD / GitOps Terraform KVM / QEMU LDAP / OAuth / RBAC Bash / Python Neutron Networking etcd Administration SCC / Pod Security StorageClass / PVC Horizontal Scaling RCA / Post-Mortems Qualifications Education Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience Experience 6–10 years of overall IT/infrastructure experience with at least 4+ years hands-on with Red Hat OCP and/or OSP Demonstrated experience managing production-grade OpenShift clusters of 20+ nodes Prior experience in financial services, banking, healthcare, or regulated enterprise environments preferred Proven track record of resolving critical P1/P2 platform incidents under time pressure Certifications (Preferred / Mandatory) Red Hat Certified Specialist in OpenShift Administration (EX280) — Preferred Red Hat Certified Engineer (RHCE) — Preferred Red Hat Certified System Administrator (RHCSA) — Mandatory Red Hat Certified Specialist in OpenStack (EX210) — Good to have CKA (Certified Kubernetes Administrator) — Good to have Good to Have Experience with OpenShift Virtualization (formerly KubeVirt) for VM workload migration Familiarity with OpenShift Service Mesh (Istio), Kiali, and Jaeger for microservices observability Exposure to multi-cluster management via Red Hat Advanced Cluster Management (RHACM) Experience with OpenShift Data Science or AI/ML workload deployment on OCP Knowledge of container image security: Quay, Cosign, Clair vulnerability scanning Understanding of compliance frameworks — PCI DSS, SOC 2, HIPAA — in containerized environments Exposure to AWS, Azure, or GCP for hybrid/multi-cloud OCP deployments Soft Skills & Working Style Strong written and verbal communication with ability to explain complex technical issues to non-technical stakeholders Self-starter who thrives in fast-paced, client-facing delivery environments Collaborative team player with cross-functional coordination skills across Dev, QA, Security, and Network teams Detail-oriented with a structured approach to documentation, change control, and knowledge transfer Willingness to participate in after-hours support windows and on-call rotations when required
More at NexTurn