Padmi

AI Ops Engineer

Bangalore · ChennaiPosted 2 months ago
Software engineeringSenior
Apply at Mastech Digital

Opens the source posting on naukri.com

Source description

About the role

View original

AIOps Engineer DevOps Artificial Intelligence Site Reliability Engineering Function Platform Engineering & AI Operations Experience 4 – 8 Years Python Kubernetes / EKS Prometheus + Grafana ML Ops Terraform / IaC CI/CD About the Role We are looking for a seasoned AIOps Engineer who sits at the intersection of platform reliability, AI/ML operations, and intelligent automation. You will design and operate AI-driven observability pipelines, build self-healing infrastructure, and embed machine-learning models into DevOps toolchains — turning telemetry noise into actionable intelligence at scale. As a trusted member of the Platform Engineering team, you will own the full lifecycle of AIOps tooling: from anomaly-detection models and intelligent alerting to predictive capacity management and LLM-assisted incident response — all running on cloud-native, containerised infrastructure. Key Responsibilities AIOps & Intelligent Automation Design and operate AI/ML-powered observability platforms (anomaly detection, root-cause analysis, predictive alerting) using tools such as Dynatrace, New Relic, Moogsoft, or open-source equivalents. Build and maintain Python-based ML pipelines (scikit-learn, PyTorch, or TensorFlow) for log analytics, metric forecasting, and event correlation. Develop intelligent auto-remediation runbooks triggered by model-detected incidents, reducing MTTR by 40%+. Integrate LLM-assisted copilots (OpenAI / Azure OpenAI) into incident management workflows for automated triage and war-room summaries. DevOps & Platform Engineering Own CI/CD pipeline design and maintenance using GitHub Actions, GitLab CI, or Jenkins; enforce shift-left security with SAST/DAST, SBOM, and Trivy scans. Provision, scale, and manage Kubernetes clusters (EKS / AKS / GKE) using Helm, Kustomize, and ArgoCD with GitOps workflows. Manage Infrastructure-as-Code using Terraform and Ansible across multi-cloud environments (AWS primary, Azure secondary). Champion containerisation best practices: image hardening, multi-stage builds, and supply-chain security with Cosign/Sigstore. Site Reliability Engineering Define, implement, and operationalise SLOs, SLIs, and error-budget burn policies aligned to business objectives. Architect and maintain a four-layer observability stack: metrics (Prometheus/Thanos), logs (OpenSearch/Loki), traces (Jaeger/Tempo), and dashboards (Grafana). Lead blameless post-mortem processes and own the incident management lifecycle end-to-end. Drive capacity planning using ML-based demand forecasting models integrated with KEDA autoscaling policies. Python Engineering & Tooling Write production-grade Python automation: alerting integrations, CMDB reconciliation, drift-detection daemons, and Slack/PagerDuty bots. Build internal CLI tools and REST APIs (FastAPI / Flask) to expose AIOps insights to engineering teams. Maintain data pipelines feeding telemetry into feature stores and model retraining loops. Required Qualifications Experience & Education 4–8 years of hands-on experience in DevOps, SRE, or Platform Engineering roles. 2+ years in an AIOps or MLOps capacity — deploying and monitoring ML models in production. Bachelor's or Master's degree in Computer Science, Information Technology, or equivalent. Technical Skills — Must Have Python (3.x): proficient in scripting, automation, REST APIs, ML libraries, and unit testing (pytest). Kubernetes: workload design, RBAC, HPA/KEDA, network policies, and multi-tenancy. Observability stack: Prometheus, Grafana, Alertmanager, OpenTelemetry, and distributed tracing. CI/CD: GitHub Actions or GitLab CI — pipeline authoring, environment promotion, and gate policies. IaC: Terraform (modules, state management, remote backends) and Helm chart authoring. Cloud: Strong hands-on Solution Architect experience on AWS and Azure — production-grade infrastructure design, multi-cloud networking, security, and cost governance. Incident Management: PagerDuty, Opsgenie, or VictorOps integrated with AIOps runbooks. Technical Skills — Strong Advantage ML frameworks: scikit-learn, Prophet, or PyTorch for time-series anomaly detection. Kafka / NATS JetStream for high-throughput telemetry streaming. Service mesh: Istio or Linkerd for traffic observability and canary rollouts. Security tooling: Trivy, Falco, OPA/Gatekeeper, OWASP LLM Top 10 awareness. LLM / SLM Optimization Hands-on experience deploying and optimising Large Language Models (GPT-4o, Claude, Llama 3) and Small Language Models (Phi-3, Mistral, Gemma) in production AIOps pipelines. Model quantisation techniques: GGUF/GGML, GPTQ, AWQ — reducing inference footprint by 2–4 without significant accuracy loss. Prompt engineering and chain-of-thought optimisation for ops-domain tasks such as log summarisation, RCA generation, and runbook synthesis. Retrieval-Augmented Generation (RAG): building domain-specific vector stores (pgvector, MongoDB Atlas Vector Search, Pinecone) backed by incident knowledge bases. LLM inference serving: vLLM, Ollama, or TGI (Text Generation Inference) on Kubernetes — autoscaling with KEDA based on token throughput. Fine-tuning SLMs on internal ops datasets using LoRA / QLoRA for domain-specific intent classification (alert triage, ticket routing, anomaly labelling). LLM observability: token usage tracking, latency p99 SLOs, hallucination guardrails, and prompt injection detection using tools such as Langfuse or LangSmith. Context-window management and cost optimisation strategies: prompt compression, semantic caching, and batching via Anthropic / OpenAI batch APIs. Model routing and A/B evaluation frameworks — switching between frontier LLMs and local SLMs based on latency, cost, and task complexity signals. Behavioural Competencies Ownership mindset — you treat production systems as if they were your own product. Data-driven decision-making — you instrument first, optimise second. Collaborative problem-solver — comfortable pairing with data science, security, and product teams. Strong written communication — able to author RCAs, design docs, and architecture decision records. Growth orientation — curious about emerging AI tooling and proactive in knowledge sharing. Nice to Have Experience on high-scale consumer platforms (50M+ users, sub-200ms p99 SLOs). Contributions to open-source AIOps or observability projects. Familiarity with MCP (Model Context Protocol) agentic architectures. AWS / GCP / Azure certifications at Professional or Specialty level. Knowledge of FinOps tooling (Kubecost, AWS Cost Explorer API integration).

One address, no account. We’ll tell you when matching roles go live.

More at Mastech Digital

Related open roles

View all roles