Padmi
Arcadia logo
Arcadia

energy intelligence platform · utility bill management

Staff Site Reliability Engineer

ChennaiPosted 3 months ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at Arcadia

Opens the source posting on shine.com

Source description

About the role

View original

As a Staff Site Reliability Engineer at Arcadia, you will be part of the SRE/Platform Engineering team in India, playing a crucial role in technical leadership. Your responsibilities will include owning and delivering SRE projects end-to-end, serving as a technical anchor for the India SRE team, designing and implementing infrastructure solutions across AWS, leading Kubernetes operations, evolving CI/CD pipelines, driving observability stack enhancements, managing database reliability, strengthening security posture, troubleshooting complex production issues, and writing essential documentation for the team. Key Responsibilities: - Own and deliver SRE projects end-to-end, including scoping, design, implementation, testing, rollout, and documentation - Serve as a technical anchor for the India SRE team, conducting design reviews, pair debugging, and mentoring engineers - Design and implement infrastructure solutions across AWS using Terraform and CloudFormation - Lead Kubernetes operations, including cluster upgrades, capacity planning, workload scaling, and GitOps deployments - Evolve CI/CD pipelines across various tools, focusing on reducing manual steps and improving deployment reliability - Drive observability stack enhancements, delivering necessary infrastructure and architectural direction for engineering teams - Manage database reliability across MySQL and PostgreSQL, including backup validation and performance tuning - Strengthen security posture through various security measures and best practices - Troubleshoot complex production issues spanning multiple domains and create runbooks for automation - Write essential documentation such as architectural decision records, operational runbooks, and troubleshooting guides - Collaborate daily with US-based SRE leadership on incident reviews, migration planning, roadmap execution, and platform strategy Qualifications Required: Must-haves: - 814 years of experience in SRE/DevOps/Cloud Engineering - Deep expertise with AWS, Terraform, Kubernetes, CI/CD pipelines, observability stack, mentorship, communication skills, automation mindset, and incident management Nice-to-haves: - Experience with FinOps practices, secrets management platforms, event-driven architectures, AI-enabled tooling, data warehouses, workflow automation platforms, and industry certifications - Exposure to working in a company that has grown through acquisitions and consolidated infrastructure environments By joining Arcadia, you will enjoy competitive compensation, a hybrid work model, comprehensive benefits including medical insurance, flexible leave policy, office in a prime location, awards, bonus, and a supportive engineering culture that values diversity, empathy, teamwork, trust, and efficiency. Arcadia is dedicated to providing equal employment opportunities and fostering a clean energy future through diversity and inclusivity. As a Staff Site Reliability Engineer at Arcadia, you will be part of the SRE/Platform Engineering team in India, playing a crucial role in technical leadership. Your responsibilities will include owning and delivering SRE projects end-to-end, serving as a technical anchor for the India SRE team, designing and implementing infrastructure solutions across AWS, leading Kubernetes operations, evolving CI/CD pipelines, driving observability stack enhancements, managing database reliability, strengthening security posture, troubleshooting complex production issues, and writing essential documentation for the team. Key Responsibilities: - Own and deliver SRE projects end-to-end, including scoping, design, implementation, testing, rollout, and documentation - Serve as a technical anchor for the India SRE team, conducting design reviews, pair debugging, and mentoring engineers - Design and implement infrastructure solutions across AWS using Terraform and CloudFormation - Lead Kubernetes operations, including cluster upgrades, capacity planning, workload scaling, and GitOps deployments - Evolve CI/CD pipelines across various tools, focusing on reducing manual steps and improving deployment reliability - Drive observability stack enhancements, delivering necessary infrastructure and architectural direction for engineering teams - Manage database reliability across MySQL and PostgreSQL, including backup validation and performance tuning - Strengthen security posture through various security measures and best practices - Troubleshoot complex production issues spanning multiple domains and create runbooks for automation - Write essential documentation such as architectural decision records, operational runbooks, and troubleshooting guides - Collaborate daily with US-based SRE leadership on incident reviews, migration planning, roadmap execution, and platform strategy Qualifications Required: Must-haves: - 814 years of experience in SRE/DevOps/Cloud Engineering - Deep expertise with AWS, Terraform, Kubernetes, CI/CD pipelines, observabili

One address, no account. We’ll tell you when matching roles go live.

More at Arcadia

Related open roles

View all roles