Source description
About the role
Cloud Operations Lead SRE / DevOps / Platform EngineeringExperience912 Years ShiftOverlap with US & EU Business Hours Role SummaryWe are seeking an experienced Cloud Operations Lead with a strong background in Site Reliability Engineering (SRE), DevOps, and Platform Engineering. The ideal candidate will be responsible for ensuring the reliability, security, and operational excellence of cloud-based platforms and services while leading a small team of engineers. This is a hands-on role with approximately 80% focus on Cloud Operations, Production Support, Reliability, and Platform Ownership, combined with leadership responsibilities. Key ResponsibilitiesLead cloud operations and production support activities across AWS-based platforms.Manage and troubleshoot Linux systems, cloud infrastructure, networking, and Kubernetes environments.Drive operational excellence through monitoring, observability, automation, and incident management.Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and Helm.Support and optimize CI/CD pipelines using GitHub Actions, Jenkins, and deployment automation tools.Design and implement monitoring, alerting, dashboards, runbooks, and operational standards.Lead vulnerability remediation, secrets management, access governance, and platform hardening initiatives.Automate infrastructure provisioning, OS/AMI upgrades, and day-2 operational activities.Support production deployments, release management, and change control processes.Collaborate with engineering teams on onboarding, platform readiness, access management, and operational best practices.Mentor and guide junior engineers while driving continuous service improvement.Required Skills (Non-Negotiable)Strong Linux Administration and TroubleshootingAWS Cloud Operations (IAM, EC2, Networking, EKS)Kubernetes Administration and Production SupportTerraform and Infrastructure as CodeCI/CD Tools (GitHub Actions, Jenkins)Monitoring & Observability (Datadog, Prometheus, Grafana, SignalFx, Nagios, or similar)Incident Management, Root Cause Analysis, and Production SupportSecurity Operations including vulnerability remediation, access management, and secrets rotationExperience working in enterprise environments with formal change management processesPreferred SkillsDNS, Proxy, Edge Services, and Networking PlatformsTeleport, Bastion Hosts, Service Accounts, and Access Management SolutionsContainer Security and Supply Chain SecurityAMI/Image Lifecycle ManagementAI-enabled Operations, Custom Agentic AI, or Hyperscaler AI ServicesLeadership ExpectationsLead a team of cloud/platform engineers.Drive operational governance, service reliability, and process standardization.Promote automation-first and reliability-first engineering practices.Partner with stakeholders across Cloud, Infrastructure, Security, and Application teams.Nice to HaveExperience in SRE, Platform Engineering, or Managed Services environments.Exposure to AI-powered operations, observability, or automation solutions.Experience supporting large-scale distributed systems and cloud-native applications. .
More at NewVison
Related open roles
DevOps / Cloud Infrastructure Engineer (AWS & Kubernetes)
Mumbai
Kafka & Kubernetes - Cloud DevOps
Mumbai
Data Center Network Engineer
Mumbai
Network Engineer - Campus, Branch And Software-Defined Wide Area Network
Mumbai
Cloud Infrastructure (Pune)
Mumbai
Cloud Operations Lead SRE DevOps Platform Engineering
Mumbai