Padmi

Remote Observability / Telemetry / SRE Engineer

Remote · United StatesPosted 1 month ago
Software engineeringUnspecified
Apply at 3B Staffing

Opens the source posting on 3bstaffing.com

Source description

About the role

View original

Job Title: Observability / Telemetry / SRE Engineer

Primary Location: Primarily Remote, Candidates Ideally Reside in Chicagoland Area

Position Type: Contract, 6 months with possible extension

Must Have:

Datadog

Prometheus

OpenTelemetry

Overview

Looking for an Observability / Telemetry / SRE Engineer to support our client's cloud modernization and acquisition-integration initiatives. This senior-level consultant will assess whether services can be effectively monitored, operated, dual-run, migrated, cut over, and rolled back using the organization's supported observability and incident-management standards.

The engineer will evaluate existing metrics, logs, traces, dashboards, alerts, service-health signals, retention requirements, PagerDuty integrations, service-level objectives, and operational runbooks. This person will identify gaps between current tooling and target-platform capabilities, define the minimum telemetry required for safe migration, and estimate the remediation effort needed to meet operational acceptance criteria.

What You Bring to the Role. (Ideal Experience)

Deep experience supporting production observability, telemetry, reliability engineering, and SRE practices.

Strong hands-on experience with:

Datadog

Prometheus

OpenTelemetry

Structured application logging

Distributed tracing

Log-ingestion and processing pipelines

Experience migrating or translating telemetry from platforms such as Datadog or Google Cloud Logging into target observability services.

Strong knowledge of metric naming, label and tag design, and cardinality management.

Experience designing, translating, and validating dashboards and alerting rules.

Knowledge of telemetry storage, retention requirements, and ingestion architecture.

Experience with PagerDuty, incident routing, escalation policies, and on-call operations.

Strong understanding of:

Service Level Indicators

Service Level Objectives

Error budgets

Service-health signals

Operational readiness criteria

Experience developing and evaluating runbooks, incident procedures, and operational support models.

Ability to distinguish cosmetic dashboard differences from an actual loss of operational visibility.

Ability to define the minimum observability and rollback signals required for a safe migration.

Understanding of instrumentation changes, telemetry schemas, service metadata, sensitive-data handling, and tenant-correlation risks.

Strong written communication skills with the ability to translate technical findings into concise engineering assessments.

What You'll Do. (Skills Used in this Position)

  • Assess whether acquisition and modernization services can be safely monitored, operated, dual-run, cut over, and rolled back.

  • Analyze existing:

  • Metrics

  • Logs

  • Traces

  • Dashboards

  • Alerts

  • Retention policies

  • Service-health indicators

  • PagerDuty integrations

  • SLOs

  • Runbooks

  • Compare current observability capabilities with the organization's target-platform standards.

  • Evaluate the effort required to migrate telemetry from existing monitoring and logging platforms.

  • Identify telemetry gaps that could affect migration safety, production support, incident response, or rollback.

  • Define minimum cutover, rollback, and service-health signals.

  • Assess metric schemas, label design, cardinality, trace propagation, structured logging, and telemetry correlation.

  • Evaluate instrumentation changes required at the application, platform, and service levels.

  • Identify sensitive-data, data-retention, tenant-isolation, and tenant-correlation concerns.

  • Perform bounded, non-production signal-parity validation when needed.

  • Partner with internal observability, platform, security, application, and incident-management teams.

  • Interpret target ingestion schemas, metadata requirements, retention standards, and operational acceptance criteria.

  • Produce concise Engineering Assessment Inputs documenting:

  • Telemetry gaps

  • Operational risks

  • Required platform dependencies

  • Instrumentation changes

  • Minimum migration-safety signals

  • Cutover and rollback requirements

  • Likely future remediation effort

One address, no account. We’ll tell you when matching roles go live.

More at 3B Staffing

Related open roles

View all roles
Remote Observability / Telemetry / SRE Engineer at 3B Staffing · Padmi