Padmi
Veeam Software logo
Veeam Software

data resilience and backup · data security posture management

Veeam Software - Staff Site Reliability Engineer

BangalorePosted 2 months ago
Software engineeringStaff+Full Time; Regular
Apply at Veeam Software

Opens the source posting on shine.com

Source description

About the role

View original

Job Description: About the Role: We are looking for a Staff Site Reliability Engineer, you will serve as a hands-on technical leader within the SRE team, guiding senior engineers, influencing product development teams, and ensuring the systems we operate are built to be reliable, scalable, and observable from the ground up. You will drive strategic initiatives, mentor others in the practice of SRE, and help define architectural best practices across our platform. This role is pivotal in aligning teams, enforcing high standards, and scaling SRE principles globally within Veeam. What Youll Do: Reliability Engineering & Resilience: - Act as a technical authority in your area, mentoring senior engineers and guiding design choices that improve service reliability and resilience - Lead the definition and enforcement of SLIs, SLOs, and error budgets; drive adherence across engineering teams - Collaborate with Staff peers across teams to align strategy and champion shared reliability standards and goals - Partner with development and product teams to proactively design for failure, build resilient architecture, and operationalize reliability from the start Observability & Operational Excellence: - Drive company-wide adoption of observability best practices and tooling - Ensure metrics, logs, and traces provide deep, actionable insights across systems - Lead complex incident responses, postmortems, and systemic reliability improvements - Promote and enforce a blameless culture of learning and continuous improvement Engineering at Scale: - Lead initiatives in infrastructure as code, deployment automation, and resilience testing - Influence the development and adoption of chaos engineering practices and release validation frameworks - Partner with platform and security teams to ensure production readiness Collaboration & Culture: - Work closely with your peer Staff Engineers to plan, align, and deliver against reliability goals - Provide architectural guidance and advocate for engineering rigor and consistency - Represent the SRE team in technical leadership forums and product planning discussions What Youll Bring: - 8+ years of experience in a Software Engineering or SRE role, including technical leadership - Demonstrated experience mentoring and guiding senior engineers - Deep expertise in building distributed systems on public cloud (Azure preferred) - Strong skills in programming (e.g., JS, Go, Typescript, Java, or C#) - Hands-on experience with observability tooling (e.g., Prometheus, Grafana, OpenTelemetry) - Mastery of infrastructure automation tools (Terraform, Pulumi) and container orchestration (Kubernetes) - Ability to communicate clearly across geographies and disciplines Bonus Skills: - Experience leading SRE initiatives across multiple product teams - Background in chaos engineering, incident learning, or performance and load testing - Familiarity with global compliance standards (ISO, SOC 2, GDPR, FedRAMP, CMMC) What Youll Get: - 18 paid vacation days, plus 4 extra global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares - Private medical coverage for you and up to four dependents - Life, accident, and disability insurance with enhanced coverage - Annual flexible wellbeing allowance for physical and mental wellness - Free confidential counselling and coaching via Employee Assistance Program (EAP), including legal and financial advice - Meal, fuel, and transportation benefits based on work arrangement - Daycare reimbursement and safe cab facility for eligible employees - Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, OReilly), mentoring, workshops, and learning events like our annual Global Day of Learning Job Description: About the Role: We are looking for a Staff Site Reliability Engineer, you will serve as a hands-on technical leader within the SRE team, guiding senior engineers, influencing product development teams, and ensuring the systems we operate are built to be reliable, scalable, and observable from the ground up. You will drive strategic initiatives, mentor others in the practice of SRE, and help define architectural best practices across our platform. This role is pivotal in aligning teams, enforcing high standards, and scaling SRE principles globally within Veeam. What Youll Do: Reliability Engineering & Resilience: - Act as a technical authority in your area, mentoring senior engineers and guiding design choices that improve service reliability and resilience - Lead the definition and enforcement of SLIs, SLOs, and error budgets; drive adherence across engineering teams - Collaborate with Staff peers across teams to align strategy and champion shared reliability standards and goals - Partner with development and product teams to proactively design for failure, build resilient architecture, and operationalize reliability from the start Observability & Operational Excellenc

One address, no account. We’ll tell you when matching roles go live.

More at Veeam Software

Related open roles

View all roles