Padmi

Automation SRE Engineer AIOps & Observability

BangalorePosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at ORANGEPEOPLE LLC

Opens the source posting on shine.com

Source description

About the role

View original

OP is partnering with a globally renowned leader in media entertainment and consumer experiences to secure a talented Automation SRE Engineer - AIOps & Observability. The team s mission is to eliminate reactive operations through high-fidelity telemetry and AI/ML-driven insight. As a hands-on contributor the Automation SRE designs observability frameworks drives SRE practices and develops the automation that connects signals to action - accelerating incident prevention detection and resolution across the global technology landscape. Responsibilities of Role: Observability Platform Engineering Design build and maintain enterprise observability platforms spanning the full range of telemetry signals including metrics logs distributed traces events and profiles - across applications infrastructure and network domains. Implement and operate observability tooling including LogicMonitor Datadog Grafana Prometheus OpenTelemetry Splunk AppDynamics (AppD) and similar platforms across infrastructure application and network domains. Instrument services applications and infrastructure with telemetry collectors exporters and agents using standards such as OpenTelemetry (OTel). Build and maintain dashboards for various personas SLO/SLI reports and alerting configurations that provide accurate actionable signals with minimal noise. Define and manage Service Level Objectives (SLOs) Service Level Indicators (SLIs) and error budgets for critical enterprise services spanning applications infrastructure and platform layers. Drive alert quality initiatives - tuning thresholds eliminating alert fatigue and building runbook-linked notification workflows. Data Engineering for Operations Build and operate telemetry data pipelines that ingest normalize enrich and route operational data at enterprise scale. Design data schemas and models for operational metrics events logs and traces to support analytics and AIOps use cases. Integrate observability data sources with data lakes streaming platforms and analytics tools to enable advanced operational reporting. Ensure data quality retention policies access controls and governance for operational datasets. Develop self-service data products and APIs that surface operational insights to engineering and operations teams. AIOps & Intelligent Automation Design and deploy AIOps capabilities including anomaly detection predictive alerting event correlation and noise suppression. Build auto-remediation workflows that trigger from observability signals reducing mean time to repair (MTTR) and manual operational effort. Apply ML/AI models to operational data for root cause correlation incident trend prediction and capacity forecasting. Integrate AIOps platforms (e.g. Dynatrace BigPanda ServiceNow ITOM or equivalent) with the observability stack to enable unified event management and automated triage across IT domains. Explore and apply Generative AI and LLM-based tooling to operational use cases including incident summarization intelligent alert triage and on-call assistant capabilities. Continuously evaluate and improve the accuracy and coverage of AIOps detections through feedback loops and model retraining. Site Reliability Engineering (SRE) Practices Champion SRE principles across the organization: toil reduction error budget policy capacity planning reliability reviews and a culture of blameless continuous improvement. Lead post-incident reviews (PIRs) and blameless postmortems; identify systemic improvement opportunities and track remediation actions to closure. Partner with platform and application engineering teams to embed reliability practices into CI/CD pipelines and service design. Develop and maintain runbooks operational playbooks and on-call response documentation for observability platform services. Participate in on-call rotation to support the observability platform and drive continuous improvement of operational procedures. Automation & Infrastructure as Code Develop automation scripts tools and integrations using Python Go or similar languages to reduce manual operational effort. Implement observability-as-code practices: manage dashboards alerts SLOs and monitors through version-controlled configuration (e.g. Terraform). Maintain CI/CD pipelines for continuous delivery of observability stack changes configuration updates and tooling enhancements. Contribute to shared automation libraries reusable modules and internal developer tools used across the broader Services and Platforms organization. Collaboration & Business Partnership Work closely with Network Engineering Platform Engineering Cloud Security and Application teams to align observability coverage with business priorities. Present observability metrics SLO performance and reliability trends to leadership and stakeholders in clear non-technical terms where appropriate. Actively engage in code reviews architectural discussions and knowledge-sharing sessions with globally distributed team members. Mentor O

One address, no account. We’ll tell you when matching roles go live.

More at ORANGEPEOPLE LLC

Related open roles

View all roles