Source description
About the role
About the Role We are seeking an experienced and highly motivated Senior DevOps, System administrator to design, implement, and manage secure, scalable, and highly available cloud-native infrastructure The candidate should possess deep expertise in Linux, OpenShift, Kubernates, CI/CD, containerization, cloud-native technologies, automation, and monitoring micro-service-based applications You will play a critical role in building and maintaining enterprise-grade DevOps platforms, optimizing application delivery pipelines, improving operational efficiency, and ensuring infrastructure reliability, security, and compliance Key Responsibilities Platform Engineering - Design, deploy, and manage enterprise-scale OpenShift clusters across Development, UAT, DR, and Production environments - Ensure high availability, scalability, resiliency, disaster recovery, and performance of container platforms - Implement Kubernetes/OpenShift best practices including: o Deployments o StatefulSets o DaemonSets o CronJobs o RBAC o Network Policies o Ingress Controllers o Storage Classes o Persistent Volumes o Operators - Manage container lifecycle, image repositories, and runtime security DevOps CI/CD - Design, implement, and maintain end-to-end CI/CD pipelines - Automate build, test, security scanning, deployment, and rollback processes - Integrate CI/CD tools including: o GitLab o ArgoCD (GitOps) o Tekton o Maven o Gradle o Nexus Repository o SonarQube o JUnit o Trivy o OWASP Dependency Check - Implement GitOps practices using ArgoCD - Manage branching strategies, pull requests, code reviews, release management, and version control Containerization Cloud-Native - Build and optimize Docker images following security and performance best practices - Deploy and manage large-scale microservice architectures - Optimize application startup, container resource allocation, JVM tuning, and memory utilization - Implement readiness, liveness, and startup probes - Configure Horizontal Pod Autoscaler (HPA) Infrastructure Automation - Strong expertise in Linux administration (RHEL, Ubuntu) - Develop Bash/Shell/Python scripts for automation - Automate operational tasks including: o Backup o Log rotation o Certificate renewal o Platform maintenance o Deployment automation - Manage Infrastructure as Code (IaC) using: o Helm Charts Middleware Web Servers - Configure and troubleshoot: o NGINX o Apache HTTP Server o IBM HTTP Server (IHS) - Configure: o Reverse Proxy o Load Balancing o SSL/TLS o HTTP/2 o WebSocket o CORS o Compression o Security Headers - Optimize web server performance and troubleshoot gateway, timeout, buffering, and proxy-related issues Database Caching - Configure and manage: o Redis Standalone o Redis Cluster - Monitor database connectivity and connection pools - Optimize HikariCP and JDBC connection pooling Messaging Streaming - Deploy and administer: o Apache Kafka Troubleshoot messaging bottlenecks and optimize throughput Monitoring, Logging Observability - Design centralized monitoring and logging platforms - Configure and manage: o Prometheus o Grafana o Alertmanager o ELK Stack (Elasticsearch, Logstash, Kibana) o Loki - Analyze application, infrastructure, and Kubernetes/Openshift logs - Create dashboards, alerts, and SLA monitoring Security Compliance - Implement DevSecOps best practices - Integrate: o ACS /ACM o SonarQube o OWASP Dependency Check - Manage: o Secrets o Certificates o RBAC o Network Policies o Security Contexts o Pod Security Standards - Ensure compliance with enterprise standards including: o ISO 27001 o GRC Policies Incident Operations Management - Troubleshoot production issues involving: o OpenShift o Redis o Kubernates, Grafana, Prometheous etc o Networking o DNS o Load Balancers o Storage o JVM o Databases - Perform root cause analysis (RCA) - Manage incidents, problem records, and change requests according to SLA guidelines - Plan and execute Disaster Recovery (DR) and failover activities SCM - Git - GitLab - GitHub Build Tools - Maven - Gradle Artifact Repository - Nexus Repository Security - SonarQube - Trivy - OWASP Dependency Check Monitoring - Prometheus - Grafana - Loki - ELK Stack Web Servers - NGINX - Apache HTTP Server - IBM HTTP Server Middleware - Tomcat - Spring Boot - Java - JVM Performance Tuning Databases - Oracle - PostgreSQL - MySQL Caching - Redis - Redis Cluster Messaging - Apache Kafka Infrastructure as Code - Helm Preferred Qualifications - Bachelors degree in Computer Science, Information Technology, or a related field - 4+ years of experience in DevOps, Platform Engineering - Hands-on experience with OpenShift Operators and Operator Lifecycle Manager (OLM) - Strong understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, SSL/TLS, Load Balancing, and Firewalls - Experience with enterprise-scale production environments supporting mission-critical applications Preferred Certifications - Red Hat Certified Specialist in OpenShift Administration - Certified Kubernetes Administrator (CKA) - Certified Kubernetes Application Developer (CKAD) - Kubernetes Security Specialist (CKS) - Red Hat Certified Engineer (RHCE) Soft Skills - Strong analytical and troubleshooting skills - Excellent communication and documentation abilities - Ability to work independently and collaboratively in cross-functional teams - Strong problem-solving and decision-making skills - Ability to manage multiple priorities in a fast-paced environment - Commitment to continuous learning and adoption of emerging technologies Experience and Education - 4 years BE/BE Tech/M CA Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at TRIGYN TECHNOLOGIES INC