Padmi
Analytic Partners logo
Analytic Partners

marketing mix modeling · commercial measurement

Site Reliability Engineer

United States · Hybrid$100k–$115k/yrPosted 3 months ago
InfrastructureUnspecifiedRegular Employee
Apply at Analytic Partners

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Own the Internal Developer Platform (IDP) as a product, treating engineering teams as customers and optimizing for reliability, usability, and delivery velocity.

Define and execute a platform roadmap aligned with business priorities, developer needs, and long-term scalability.

Design, build, and evolve paved roads for application delivery, including CI/CD pipelines, infrastructure templates, service scaffolding, and standardized deployment patterns.

Build self-service capabilities that enable teams to provision, deploy, observe, and operate services with minimal friction.

Create and maintain reusable platform abstractions across AWS and Azure that standardize security, reliability, networking, and observability.

Reduce developer cognitive load by abstracting unnecessary complexity while enforcing clear guardrails for security, cost, and compliance.

Partner closely with application, product, and security teams to embed reliability, scalability, and security by design.

Establish and evolve platform standards for logging, monitoring, alerting, tracing, and incident response workloads.

Define, measure, and manage SLIs, SLOs, and error budgets for shared platform services.

Drive the reduction of operational toil through automation, standardization, and platform-first solutions.

Ensure shared platform services meet high standards for availability, performance, resilience, and scalability.

Own system-to-system integration and messaging patterns used across the platform.

Lead capacity planning, demand forecasting, and performance tuning for platform services.

Plan and execute zero-downtime upgrades, migrations, and releases of platform components.

Lead platform-level incident response workflows, post-incident reviews, and drive systemic improvements rather than one-off fixes.

Evaluate incoming platform requests and translate them into scalable, productized capabilities.

Mentor engineers and drive platform adoption through documentation, enablement, and technical evangelism.

Participate in a 24x7 on-call rotation as an escalation point for platform reliability and availability issues.

Operate effectively in ambiguous problem spaces, making sound architectural and product decisions with limited guidance.

More at Analytic Partners

Related open roles

View all roles