Source description
About the role
Rol About the Role We are looking for a Full Stack Observability Engineer not a traditional DevOps or Operations professional, but a specialist who deeply understands the why behind every signal, metric, trace, and log across the entire technology stack. This role demands someone who can see the full picture: from bare-metal infrastructure and network fabric to application layers and business transactions — and turn that visibility into intelligent, predictive action. Key Responsibilities End-to-End Observability — Design, implement, and own observability coverage across infrastructure, applications, networking, and data centre environments; ensure no blind spots exist across the stack Tooling & Platform Ownership — Administer and optimize the full observability toolchain (Datadog, Prometheus, Grafana, New Relic, Dynatrace, OpsManager); drive integration across tools for unified visibility AIOps & Intelligent Monitoring — Build and operationalize ML-based models for anomaly detection, noise reduction, and event correlation; develop automation using Python and PowerShell scripting Predictive Maintenance — Leverage historical telemetry and ML models to predict failures, capacity breaches, and performance degradation before they impact the business Dashboards & Alerting — Design meaningful, stakeholder-ready dashboards and intelligent alerting strategies that reduce alert fatigue and surface actionable insights Stakeholder Communication — Collaborate with global teams, client stakeholders, and leadership; translate observability findings into clear, business-relevant language Continuous Improvement — Identify observability gaps, propose remediation roadmaps, and drive maturity uplift across monitoring disciplines Mandatory Skills Area Requirement Observability Tools Datadog, Prometheus, Grafana, New Relic, Dynatrace, OpsManager — end-to-end proficiency required Monitoring Scope Infrastructure, Application, Network, Data Centre — full stack coverage AIOps ML model integration for event correlation & anomaly detection Scripting Python and PowerShell — mandatory Predictive Maintenance Practical experience with predictive alerting and failure forecasting Communication Professional business communication for global client engagement e & responsibilities
More at Ikrux