Source description
About the role
Site Reliability Engineer (SRE) / DevOps Engineer – AWS, Kubernetes (EKS), CI/CD Location: Remote / base location Noida Experience: 7+ years (flexible based on depth in AWS + Kubernetes + production ops) Role Summary: We are looking for an experienced Site Reliability Engineer (SRE) to own reliability, scalability, automation, and operational excellence for cloud-native platforms running on AWS and Kubernetes (EKS). You will build and operate CI/CD pipelines, standardize infrastructure provisioning, implement monitoring/alerting, drive incident response, and partner with architects and engineering teams to deliver secure, cost- efficient, highly available systems. Key Responsibilities: Reliability & Operations (SRE Core) • Own availability, latency, performance, and capacity for production workloads. • Define and track SLIs/SLOs, error budgets, and reliability KPIs. • Run incident management (on-call, triage, RCA/postmortems, preventive actions). • Improve MTTR through automation, runbooks, and self-healing patterns. Kubernetes & Platform Engineering (EKS) • Operate and evolve AWS EKS clusters (multi-namespace, multi-env). • Manage deployments using Helm/Kustomize, and enable safe rollouts (blue/green, canary). • Handle ingress (ALB/NLB), service discovery, autoscaling (HPA/VPA/Cluster Autoscaler). • Implement policies using RBAC, NetworkPolicies, Pod Security, secrets management. CI/CD & Release Engineering • Build and maintain CI/CD pipelines using GitHub Actions / Jenkins / GitLab CI (based on org standard). • Enforce best practices: pipeline templates, approvals, artifact versioning, rollback strategy. • Container build & security scanning workflows (SAST/DAST/image scanning) and SBOM where required. • Promote everything in Git culture: infra/app configs + deployment manifests. Infrastructure as Code & Cloud Automation • Provision cloud infra using Terraform / CloudFormation / CDK (as applicable). • Implement reusable modules for VPC, IAM, EKS, RDS, S3, CloudFront, WAF, Route53, etc. • Manage multi-account AWS access patterns (dedicated Dev accounts, IAM roles, federation/SSO). • Implement cost controls: tagging standards, budgeting alerts, right-sizing, savings plans guidance. Observability (Monitoring, Logging, Tracing) • Own observability stack: Prometheus, Grafana, CloudWatch, Alertmanager. • Implement metrics and dashboards for K8s cluster health + application SLIs. • Implement logging pipeline (e.g., Fluent Bit/Fluentd → CloudWatch/ELK/OpenSearch). • Enable distributed tracing (OpenTelemetry / X-Ray / Jaeger) where needed. Security & Governance (DevSecOps) • Implement edge and app protection using AWS WAF / CloudFront, and coordinate rollout safely. • Enforce least privilege IAM, secrets handling, KMS encryption, secure networking. • Support vulnerability remediation (CVEs), patching processes, and audit readiness. Collaboration & Process • Partner with Engineering + Architecture to review designs and keep solutions simple. • Create/maintain runbooks, SOPs, onboarding guides, and deployment standards. • Raise infra requests and coordinate with platform/GT teams to ensure timely provisioning. Must-Have Skills • Strong hands-on experience with AWS (EKS, EC2, IAM, VPC, ALB/NLB, CloudWatch, S3, RDS). • Deep experience managing Kubernetes in production, preferably EKS. • Strong CI/CD ownership: pipelines, environments, release controls, rollback strategies. • Observability experience: Prometheus + Grafana, alerting, dashboards, incident response. • Infrastructure as Code: Terraform (preferred) or CloudFormation/CDK. • Solid Linux + networking fundamentals (DNS, TLS, load balancing, routing). • Strong scripting skills: Python/Bash. Good-to-Have Skills • Service mesh (Istio/Linkerd), OPA/Gatekeeper, Kyverno. • External SaaS monitoring (Grafana Cloud, Datadog, New Relic). • GitOps tools: Argo CD / Flux. • Experience with WAF tuning, bot rules, rate limiting, and safe production rollout. • PostgreSQL/MySQL operations basics (performance, monitoring, connection pooling). • Experience with high-volume ingestion/data pipelines is a plus. AWS,GitHub Actions , Jenkins , GitFlows, Terraforms, EKS If you have relevant exp share your resume to [HIDDEN TEXT]
More at Innodata