Source description
About the role
Senior DevOps Engineer: Company: Ethara AI Location: 5th Floor, Plot No. 273, Udyog Vihar II Road, Udyog Vihar Phase 1, Sector 20, Gurugram, Haryana 122016 Employment Type: Full-Time Experience Required: 5+ Years Role Overview We are looking for a highly skilled Senior DevOps Engineer with strong expertise in cloud infrastructure, containerization, and automation. In this role, you will be responsible for designing, deploying, and managing scalable, secure, and high-performance infrastructure that supports AI-driven platforms and production-grade systems. You will play a key role in driving DevOps maturity, implementing best practices, and enabling seamless deployment and observability across distributed systems. Key Responsibilities Design and manage scalable, secure cloud infrastructure on AWS, including VPC, IAM, EC2, S3, RDS, ECR, Route 53, and Secrets Manager Build and maintain Infrastructure as Code (IaC) using Terraform, Pulumi, or AWS CDK Develop and manage CI/CD pipelines using Git, GitHub Actions, Jenkins, and ArgoCD with GitOps practices Design, build, and optimize containerized applications using Docker , including image creation, multi-stage builds, and runtime optimization, and deploy them on Kubernetes (EKS) using Helm, HPA, RBAC, and IRSA Own end-to-end container lifecycle management , including image optimization, security hardening, versioning, and efficient deployment across environments Implement advanced Kubernetes capabilities such as Karpenter, Istio service mesh, mTLS, and canary deployments Set up and manage observability and monitoring systems using CloudWatch, Prometheus, Grafana, Loki, and OpenSearch Design and manage streaming and messaging systems using Kafka, MSK, and Schema Registry Ensure robust security practices, including IAM policies, container security, TLS/SSL, and shift-left security approaches Handle incident management, perform root cause analysis (RCA), and improve system reliability Optimize infrastructure for cost, performance, and scalability Collaborate with engineering and AI teams to support LLM infrastructure and deployment workflows, including AWS Bedrock integration Required Skills & Qualifications 5+ years of experience in DevOps / Cloud / Infrastructure Engineering Strong hands-on experience with AWS services: VPC, IAM, EC2, S3, RDS, ECR, Route 53, Secrets Manager Strong hands-on experience with Docker , including containerization, image optimization, multi-stage builds, and debugging containerized environments Deep expertise in Kubernetes (EKS), including Helm, HPA, RBAC, IRSA, Karpenter, and Istio Strong experience with CI/CD tools: Git, GitHub Actions, Jenkins, ArgoCD Hands-on experience with Infrastructure as Code tools such as Terraform (preferred), Pulumi, or AWS CDK Experience with monitoring and observability tools: CloudWatch, Prometheus, Grafana, Loki, OpenSearch Working knowledge of Kafka/MSK and distributed streaming systems Strong understanding of Linux systems, networking, DNS, TLS/SSL, ALB/NLB Proficiency in scripting (Python and/or Bash) Strong understanding of security best practices (IAM, RBAC, container security) Experience in incident management and reliability engineering Preferred / Bonus Skills Experience with AWS Bedrock and AI/ML infrastructure Familiarity with service mesh (Istio) and mTLS architectures Exposure to MLOps tools such as Kubeflow or MLflow Proven experience in cost optimization at scale Prior experience working in AI/LLM-driven environments Certifications (Preferred) AWS Solutions Architect certifications Certified Kubernetes Administrator (CKA) or Developer (CKAD) Terraform Associate Certification DevSecOps or Security+ Certification Why Join Us Work on cutting-edge AI/ML infrastructure and LLM systems Design and scale production-grade distributed systems Be part of a high-growth, high-impact engineering team Opportunity to shape DevOps architecture and best practices Competitive compensation and continuous learning opportunities