Padmi

Senior Staff Site Reliability Engineer (SRE)

IndiaPosted 2 months ago
Software engineeringStaff+Full Time; Regular
Apply at Arrow Components

Opens the source posting on shine.com

Source description

About the role

View original

Job Description We are seeking a Sr Staff Site Reliability Engineer on a longterm basis during USA hours who brings deep software engineering roots alongside SRE expertise. This individual will help shape and scale the reliability of our global cloud platform, bringing the fullstack perspective of someone who has built and shipped software and now drives reliability from the inside out. This is a Senior Stafflevel technical leadership role with organizationwide influence. You will define and drive reliability strategy across our multicloud infrastructure (AWS and GCP), establish architectural standards, and ensure our backend systems operate with exceptional availability, scalability, and resilience. You will also collaborate with strategic partners and engineering teams to enable our organization as a cloudintegrated service, leading technical discussions and ensuring secure and reliable integrations. This is a longterm position for someone who thrives at the intersection of software development and reliability engineering. The ideal candidate has handson development experience, understands the complete software delivery lifecycle, and brings an endtoend systems perspective from code commit to production operation. Responsibilities Define and drive Organizations SRE strategy across engineering teams. Establish reliability standards, architectural guardrails, and production readiness frameworks. Initiate, participate in, and review architectural changes leveraging development experience to ensure reliability and operability are built in, not bolted on. Apply SDLC knowledge to reliability decisions engage early in design and architecture reviews to embed reliability, testability, and operability as firstclass requirements. Proactively identify systemwide gaps continuously assess the platform for reliability blind spots, missing observability, or architectural debt, and drive initiatives to close them without waiting to be asked. Bridge development and SRE teams translate between engineering intent and operational reality, serving as a technical liaison who can read code, review PRs, and contribute to servicelevel design decisions. Design and maintain highly available, multiregion, multicloud systems. Ensure platform reliability supporting millions of IoT devices globally. Guide engineering teams in building faulttolerant, scalable microservices and monolithic systems. Define and enforce SLIs, SLOs, and error budgets. Lead architecture reviews and production readiness reviews. Partner with strategic teams to deliver our organization as a cloudintegrated service and support partner integrations. Improve and streamline production release processes. Implement safe deployment strategies (canary, blue/green, progressive delivery). Build CI/CD guardrails to reduce deployment risk and improve reliability. Develop and mature observability strategies across infrastructure and services. Lead highseverity incident response, facilitate blameless postmortems, and drive systemic improvements to prevent recurring issues. Qualifications 10+ years of combined software engineering and SRE/infrastructure experience, with a clear progression from development into reliability or platform engineering. Deep understanding of the complete Software Development Lifecycle (SDLC) enabling wellinformed reliability and design decisions across all phases of software delivery. Strong software development background with handson experience building and shipping production software enabling effective design collaboration, codelevel review, and reliabilitydriven architectural input. Endtoend system comprehension ability to reason about the full stack from device/client behavior through API layer, backend services, data stores, and infrastructure, connecting the dots across teams and domains. Selfdirected gap identification demonstrated initiative in spotting reliability, scalability, or process gaps and driving improvements without needing explicit direction. Collaborative crossteam communication proven ability to work across engineering, product, and operations teams; comfortable influencing without authority and presenting technical decisions to both technical and nontechnical stakeholders. Proven experience operating largescale distributed systems in production. Strong handson expertise with AWS and GCP cloud platforms. Deep experience with Kubernetes in production environments. Advanced knowledge of Terraform, including modular design and infrastructure governance. Strong understanding of distributed systems, networking, and system reliability principles. Experience supporting Javabased monolithic systems and microservices architectures. Proficiency in Python for automation and tooling. Experience with modern observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, etc.). Strong debugging, incident response, and root cause analysis skills. Security knowledge in Jo

One address, no account. We’ll tell you when matching roles go live.

More at Arrow Components

Related open roles

View all roles