Source description
About the role
Role: Datadog Observability Architect (Senior / Lead) A senior observability specialist who can architect Datadog end to end, set the standards, and mentor a cloud/CI-CD team that lacks deep Datadog expertise. Must-have skills - Deep, hands-on Datadog expertise at architect level (design, not just use) across APM, infrastructure monitoring, logs, synthetics, dashboards, and monitors - Observability fundamentals: metrics, logs, and traces, plus OpenTelemetry - SLO, SLI, and error-budget design - APM and Database Monitoring for root-cause analysis (SQL Server a solid plus) - Alerting strategy and alert-noise reduction - Log and observability pipelines: routing, filtering, and cost control - Datadog cost optimization: ingestion, indexing, retention - Datadog managed as code with Terraform or OpenTofu - AWS environment experience - Integrations with CI/CD (Harness or similar), Jira or ServiceNow, and webhooks Nice-to-have - Datadog and AWS certifications - SolarWinds-to-Datadog migration experience - Synthetic and end-to-end journey monitoring - Automation or self-healing (webhook-triggered pipelines) - Scripting: PowerShell, SQL, Python - Regulated or financial-services background Experience-Level: roughly 8+ years in SRE, observability, or DevOps, with 4+ years hands-on Datadog at scale. Architect-level, comfortable leading a small pod and client-facing Role: Datadog Observability Architect (Senior / Lead) A senior observability specialist who can architect Datadog end to end, set the standards, and mentor a cloud/CI-CD team that lacks deep Datadog expertise. Must-have skills - Deep, hands-on Datadog expertise at architect level (design, not just use) across APM, infrastructure monitoring, logs, synthetics, dashboards, and monitors - Observability fundamentals: metrics, logs, and traces, plus OpenTelemetry - SLO, SLI, and error-budget design - APM and Database Monitoring for root-cause analysis (SQL Server a solid plus) - Alerting strategy and alert-noise reduction - Log and observability pipelines: routing, filtering, and cost control - Datadog cost optimization: ingestion, indexing, retention - Datadog managed as code with Terraform or OpenTofu - AWS environment experience - Integrations with CI/CD (Harness or similar), Jira or ServiceNow, and webhooks Nice-to-have - Datadog and AWS certifications - SolarWinds-to-Datadog migration experience - Synthetic and end-to-end journey monitoring - Automation or self-healing (webhook-triggered pipelines) - Scripting: PowerShell, SQL, Python - Regulated or financial-services background Experience-Level: roughly 8+ years in SRE, observability, or DevOps, with 4+ years hands-on Datadog at scale. Architect-level, comfortable leading a small pod and client-facing