Source description
About the role
Engineers here own services end-to-end—from design to production reliability.
Important: This is not a system administrator role. We are explicitly hiring an engineering leader in reliability.Engineering degree is an absolute requirement (BS/MS in CS/CE/EE or closely related engineering field).
What You’ll Do
- Own reliability outcomes for critical services: availability, latency, incident rate, and recovery time.
- Design and build reliable, scalable distributed systems that support mission-critical healthcare workflows.
- Define and operationalize SLOs/SLIs and error budgets; drive adoption across teams and use them to prioritize work.
- Lead incident response for high-severity issues; improve on-call effectiveness and reduce alert fatigue.
- Run blameless postmortems and ensure follow-ups are implemented, measured, and stick.
- Write software to eliminate operational toil: automation, self-service tooling, guardrails, and developer platforms.
- Raise the bar on observability (metrics/logs/traces), alerting strategy, and operational readiness.
- Improve resilience through capacity planning, load testing, performance tuning, and failure testing.
- Mentor engineers (SRE and product engineers) on reliability practices, debugging, and production ownership.
- Drive cross-team improvements like production readiness reviews, release safety (progressive delivery), and standard runbooks.
What We’re Looking For
-
Required
-
Engineering degree is mandatory: BS/MS in Computer Science, Computer Engineering, Electrical Engineering, or a closely related engineering field.
-
6+ years experience in software engineering, SRE, infrastructure/platform engineering, or related.
-
Strong programming skills in Go, Python, Java, or similar (production-quality code).
-
Proven experience building and operating production backend services or distributed systems.
-
Meaningful experience in on-call rotations, incident leadership, and post-incident improvement execution.
-
Strong debugging ability across complex systems: latency, saturation, cascading failures, dependency issues.
-
Experience with cloud infrastructure (AWS preferred, GCP/Azure acceptable).
-
Strong Signal
-
You’ve owned reliability for customer-facing services with clear, measurable improvements (e.g., higher availability, lower MTTR).
-
You’ve built internal platforms/tooling that made other engineers faster and reduced operational burden.
-
You’ve worked in an SRE culture with SLOs, error budgets, and blameless postmortems.
-
You’ve led multi-quarter reliability initiatives spanning multiple teams/services.
-
Technologies We Work With (Examples)
-
Cloud: AWS
-
Containers: Docker, Kubernetes
-
Infrastructure as Code: Terraform
-
Observability: Prometheus, Grafana, OpenTelemetry
-
Languages: Go, Python, TypeScript
-
CI/CD: GitHub Actions
(Experience with everything isn’t required—strong fundamentals and learning velocity matter most.)
What This Role Is Not
To be explicit, this role is not
- System administration / IT ops / helpdesk
- Manual server patching as a primary responsibility
- A “click-ops” cloud operator role
This is a senior engineering role focused on software-driven reliability and platform engineering.
Why Join PBN
- Build and operate mission-critical healthcare infrastructure that supports real patient workflows.
- High impact: reliability work directly improves customer trust and revenue-critical operations.
- Small team with high ownership, autonomy, and ability to influence architecture.
- Strong engineering culture focused on automation, simplicity, and measurable outcomes.
Compensation
The base pay range for this role is $120,000 – $150,000 per year.
Ready to apply?
Powered by
First name *
Last name *
Email *
LinkedIn URL
Resume *
Click to upload or drag and drop here
Are you willing to undergo a background check, in accordance with local law/regulations? *
Yes
No
Are you willing to undergo a background check, in accordance with local law/regulations? *
Yes
No
Are you legally authorized to work in the United States? *
Yes
No
Will you now, or in the future, require sponsorship for employment visa status (e.g. H-1B visa status)? *
Yes
No
Are you comfortable working in a remote setting? *
Yes
No
The base salary range for this role is $120k – $150k annually. Is this range aligned with your expectations? *
Yes
No
Apply
Req ID: R18
More at Practice by Numbers
Related open roles
Sr. FullStack Software Engineer
Seattle · Hybrid
Senior Software Engineer – Core VoIP / SIP Backbone
Delhi NCR · Onsite
Technical Leader – Engineering
Delhi NCR · Onsite
Sr. Site Reliability Engineer
Seattle · Hybrid
Senior Product Manager (Technical)
Remote · Seattle
Senior Software Engineer - Full Stack - Payments
Delhi NCR · Onsite
