Padmi

Site Reliability Engineer

MumbaiPosted 1 month ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at Zlendo Technologies

Opens the source posting on shine.com

Source description

About the role

View original

Key Responsibilities: Kubernetes Mastery: Design, manage, and optimize GKE (Google Kubernetes Engine) clusters. Act as the subject matter expert for all K8s-related tasks, including resource scaling, networking, and security.ELK Stack Administration: Full ownership of the ELK (Elasticsearch, Logstash, Kibana) infrastructure. This includes:Onboarding new applications and log sources.Managing User Access (RBAC) and security roles.Creating advanced Kibana dashboards and alerting systems.Data Tier Management:Redis: Deploy and manage Redis Sentinel for high availability; handle instance creation and performance tuning.Apache Kafka: Manage Kafka clusters, including topic creation, replication factor management, and partition balancing.Automation & POCs: Drive innovation by performing POCs for multi-tier applications. Build custom automation to reduce manual toil across the entire stack.Cloud Infrastructure: Manage GCP resources using Infrastructure as Code (Terraform/Ansible) with a focus on cost-efficiency and 99.99% availability. Required Technical Skills: Orchestration: Expert-level knowledge of Kubernetes (GKE) and Docker.Observability: Deep experience in ELK Stack administration (not just searching logs, but managing the cluster health and user permissions).Messaging & Caching: Hands-on experience managing Apache Kafka (Topics/Replication) and Redis (Sentinel/Clustering).Automation: Proficiency in Python or Java for building automation tools and conducting complex technical POCs.CI/CD: Experience with ArgoCD, Jenkins, or GitLab CI/CD for automated application delivery.Cloud: Strong knowledge of GCP (VPC, IAM, GKE, Cloud Storage). Key Responsibilities: Kubernetes Mastery: Design, manage, and optimize GKE (Google Kubernetes Engine) clusters. Act as the subject matter expert for all K8s-related tasks, including resource scaling, networking, and security.ELK Stack Administration: Full ownership of the ELK (Elasticsearch, Logstash, Kibana) infrastructure. This includes:Onboarding new applications and log sources.Managing User Access (RBAC) and security roles.Creating advanced Kibana dashboards and alerting systems.Data Tier Management:Redis: Deploy and manage Redis Sentinel for high availability; handle instance creation and performance tuning.Apache Kafka: Manage Kafka clusters, including topic creation, replication factor management, and partition balancing.Automation & POCs: Drive innovation by performing POCs for multi-tier applications. Build custom automation to reduce manual toil across the entire stack.Cloud Infrastructure: Manage GCP resources using Infrastructure as Code (Terraform/Ansible) with a focus on cost-efficiency and 99.99% availability. Required Technical Skills: Orchestration: Expert-level knowledge of Kubernetes (GKE) and Docker.Observability: Deep experience in ELK Stack administration (not just searching logs, but managing the cluster health and user permissions).Messaging & Caching: Hands-on experience managing Apache Kafka (Topics/Replication) and Redis (Sentinel/Clustering).Automation: Proficiency in Python or Java for building automation tools and conducting complex technical POCs.CI/CD: Experience with ArgoCD, Jenkins, or GitLab CI/CD for automated application delivery.Cloud: Strong knowledge of GCP (VPC, IAM, GKE, Cloud Storage).

One address, no account. We’ll tell you when matching roles go live.

More at Zlendo Technologies

Related open roles

View all roles