Padmi

Senior Site Reliability engineer -SRE - Platform Engineers

Hyderabad · Delhi NCR · ChennaiPosted 2 months ago
Infrastructure And DatabasesSenior
Apply at Hucon Solutions

Opens the source posting on naukri.com

Source description

About the role

View original

Job Description Site Reliability Engineer (SRE) Specialization: Kubernetes | Google Cloud Platform (GCP) Role Overview We are looking for a highly skilled Site Reliability Engineer (SRE) to build, operate, and continuously improve highly available, scalable, and observable platforms running on baremetal Kubernetes clusters and Google Kubernetes Engine (GKE). The ideal candidate brings deep Kubernetes expertise, strong cloud-native experience on GCP, and a passion for reliability, automation, and operational excellence. This role works closely with application, platform, and architecture teams to ensure production systems are resilient, secure, and performant at scale. Experience Range - 4 to 12 years Locations: Chennai, Hyderabad, Noida and Gurgaon only Key Responsibilities • Design, operate, and support Kubernetes platforms across baremetal clusters and GKE • Ensure high availability, scalability, performance, and reliability of production systems • Implement and manage GitOps-based deployment workflows using tools like Argo CD • Build, maintain, and optimize CI/CD pipelines using tools such as GitHub Actions, Harness, CircleCI, or equivalent • Deploy and manage applications using Helm, including canary and progressive delivery strategies • Implement comprehensive observability using Prometheus, Grafana, Loki, and Tempo • Proactively monitor systems, troubleshoot incidents, and perform root cause analysis (RCA) • Partner with development teams to improve service reliability, scalability, and operational maturity • Provision and manage cloud infrastructure on Google Cloud Platform (GCP) • Automate infrastructure and platform operations using Infrastructure as Code (IaC) and scripting • Drive continuous improvements in resilience, automation, and operational efficiency Required Skills & Qualifications • Strong hands-on experience with Kubernetes architecture and administration • Experience managing both bare-metal Kubernetes clusters and Google Kubernetes Engine (GKE) • Solid understanding of Google Cloud Platform (GCP) services and networking concepts • Proven experience with GitOps practices and tools such as Argo CD • Proficiency with CI/CD tools (GitHub Actions, Harness, CircleCI, or similar) Practical experience with: Helm Canary / progressive deployments • Strong expertise in observability and monitoring: Prometheus Grafana Loki Tempo • Experience with Terraform for infrastructure provisioning • Understanding of modern API technologies such as GraphQL • Familiarity with API management platforms (Apigee Edge, Apigee X) • Knowledge of CDN and edge services (e.g., Akamai) Good to Have • Working knowledge of Java (Spring Boot) and/or Node.js framework • Understanding of microservices architecture and service-to-service communication • Experience with Ansible or similar configuration management tools • Exposure to hybrid or multicloud environments • Experience in performance tuning and cost optimization on GCP • Understanding of Kubernetes and cloud security best practices • SRE experience aligned with SLIs, SLOs, and error budgets Soft Skills • Strong analytical and troubleshooting skills • Ownership mindset with a focus on automation and reliability • Ability to work effectively in a fast-paced, collaborative environment • Clear communication skills with cross-functional stakeholders Experience • 4–12 years of experience in SRE, DevOps, or Platform Engineering roles

One address, no account. We’ll tell you when matching roles go live.

More at Hucon Solutions

Related open roles

View all roles