Padmi

Site Reliability Engineer

IndiaPosted 2 months ago
Software engineeringSeniorFull Time
Apply at aviato consulting

Opens the source posting on foundit.in

Source description

About the role

View original

Aviato Consulting is seeking an experienced Site Reliability Engineer to join our growing team. This isn't just another SRE role; it's an opportunity to own critical infrastructure , drive technical strategy , and shape the reliability culture for major Australian and EU clients, all within a supportive, G-inspired environment built on transparency and collaboration. What's In It For You Learn from the Best: Report directly to and receive mentorship from our Head of SRE, an experienced ex-Google Manager . High-Impact Projects: Take ownership of complex GCP environments for diverse, significant clients across Australia and the EU. Drive Innovation, Not Just Tickets: Architect solutions and implement cutting-edge practices like SLOs, error budgets, and predictive AI anomaly detection to proactively improve systems. A Culture That Works: Founded by ex-Googlers, we foster a transparent, collaborative, and low-bureaucracy environment. What You'll Do (Your Impact): Own & Architect Reliability: Design, implement, and manage highly available, scalable architectures on Google Cloud Platform (GCP). Master Kubernetes & AI Infrastructure: Architect and manage production-grade Kubernetes clusters (GKE), and support high-performance infrastructure for AI/ML workloads (e.g., GPU/TPU pools). Drive Automation & IaC: Lead robust automation strategies using Terraform, Ansible, and scripting (Python, Go, Bash) for CI/CD pipelines. Elevate Observability with AIOps: Architect monitoring, logging, and alerting using Grafana, Dynatrace, and Sentry, integrating AI-driven insights for automated root-cause analysis. Lead Incident Response: Spearhead incident management, conduct blameless post-mortems, and leverage GenAI to automate runbook generation. Champion SRE Principles: Actively promote SLOs, SLIs, and error budgets, and mentor team members. What You'll Bring (Your Expertise): Proven SRE Experience: 5+ years of hands-on experience in a Site Reliability Engineering or Cloud Engineering role focusing on production systems. Deep GCP & Kubernetes Knowledge: Demonstrable expertise in core GCP services and managing Kubernetes clusters in production (GKE highly desirable). Infrastructure as Code Mastery: Significant experience using Terraform in complex environments. Automation & Scripting Prowess: Strong proficiency in Python or Go for automating operational tasks. Next-Gen Observability Expertise: Experience with modern monitoring tools, preferably with exposure to AI-driven alerting and logs clustering. Problem-Solving Acumen: Strong analytical skills with experience leading incident response for critical systems. (Desirable): Experience with AI infrastructure (vector databases, LLM inference), or API Management platforms like Apigee. Technologies We Use (You'll Master): Cloud: Google Cloud Platform (GCP) Containerisation & Orchestration: Kubernetes (GKE), Docker Infrastructure & Automation: Terraform, Ansible Monitoring & AIOps: Grafana, Dynatrace, Sentry, Google Cloud Operations Suite CI/CD: Jenkins, GitHub Actions, Bamboo (or similar) Scripting & AI Tooling: Python, Go, Bash, GitHub Copilot, Gemini/Vertex AI Collaboration: JIRA, Confluence, Slack Work Timings : 5am to 2 pm IST Ready to Elevate Your SRE Career If you're a passionate Senior SRE ready to tackle complex challenges on GCP, work with leading clients, and benefit from exceptional mentorship in a fantastic culture, Aviato is the place for you. Apply now and help us build the future of reliable cloud infrastructure!

One address, no account. We’ll tell you when matching roles go live.

More at aviato consulting

Related open roles

View all roles