Padmi
InvestorFlow logo
InvestorFlow

private equity CRM · deal flow management

Site Reliability Engineer II

DO · OnsitePosted 5 months ago
InfrastructureSeniorFull Time
Apply at InvestorFlow

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Design and implement comprehensive monitoring strategies rather than owning observability platforms outright.

Collaborate with DevOps and Engineering on shared observability platforms (Grafana, Prometheus/Loki, Azure Monitor/Application Insights).

Define golden signals dashboards, measure SLOs/SLIs/error budgets, and help implement actionable alerting.

Drive structured logging standards, distributed tracing patterns, and OpenTelemetry implementation standards for teams to deploy and SRE to validate.

Conduct monitoring/auditing of production systems to ensure instrumentation completeness.

Take ownership of production incident response, lead incident handling, and drive remediation.

Conduct blameless post-incident reviews and ensure follow-through on action items.

Continuously improve operational processes, reliability practices, and team readiness.

Monitor system resource utilization and forecast future needs.

Tune autoscaling configurations in partnership with Engineering teams.

Evaluate capacity efficiency and support cost optimization strategies.

Validate DR environments and test failover processes—not build them.

Ensure DR capabilities are functioning as-designed with clear documentation.

Define and lead regular DR drills in partnership with Engineering/Platform teams.

Work with the Non-Functional Testing team on resilience and DR scenario simulations.

Support chaos experiment planning and validation as a nice-to-have capability.

More at InvestorFlow

Related open roles

View all roles