Source description
About the role
Client: TCS | Engagement: Full-time | Work mode: ONSITE | Rate: 0-85 USD | Experience: 10+ | Publisher job id: JOB-000135
Role Overview
The role is for a Senior AIOps ML Engineer responsible for designing, building, and optimizing a Lakehouse architecture for multi-domain observability data at petabyte scale. This includes developing robust data ingestion and transformation pipelines, implementing advanced machine learning models for AIOps (anomaly detection, root-cause analysis, incident forecasting), and managing the full MLOps lifecycle. The engineer will also focus on platform productionization, security, compliance, and fostering engineering standards. Key Responsibilities: Design and evolve the Lakehouse schema (Delta Lake / Apache Iceberg) for multi-domain observability data at petabyte scale. Build and maintain robust ingestion pipelines from the OTel Collector through Kafka to the Lakehouse, ensuring exactly-once semantics and strict schema enforcement. Implement dbt transformation models to generate mart-ready, denormalized fact and dimension tables for each of the six domains. Define and enforce data quality contracts, establishing SLAs for data freshness, completeness, and cardinality budgets per mart. Optimize query performance utilizing partitioning strategies, Z-ordering, bloom filters, and materialized views tailored for time-series patterns. Design, train, and deploy machine learning models for streaming multivariate anomaly detection, root-cause analysis, and incident forecasting across all six mart domains. Build low-latency streaming inference pipelines (Flink / Spark Streaming) for real-time anomaly scoring on APM, infrastructure, and security signals. Develop sophisticated log intelligence models—including clustering (DRAIN3 / LogBERT), NLP classification, and error deduplication—over the Log mart. Implement unsupervised and semi-supervised methods for User Experience frustration detection and KPI correlation analysis. Own the ML feature store, managing feature engineering, versioning, backfill pipelines, and point-in-time correct joins for training datasets. Instrument model performance tracking, including drift detection, accuracy monitoring, and automated retraining triggers. Design and operate the end-to-end AIOps workflow, spanning signal ingestion, feature computation, model inference, alert routing, and auto-remediation hooks. Build high-performance model serving infrastructure—supporting real-time REST/gRPC endpoints and async batch scoring—with strict p99 latency SLOs. Integrate AIOps insights with incident management platforms (PagerDuty, Opsgenie) and internal runbooks to deliver enriched, noise-reduced alerting. Define and publish metrics from the Business KPI mart to quantify the blast radius, revenue loss, and affected user counts for each incident. Partner with the Security team to build the Security mart schema, including threat feed ingestion, UEBA baselines, and CVE correlation pipelines. Train anomalous-access and lateral-movement detection models, tuning precision/recall thresholds in collaboration with the SOC team. Ensure all data handling across the marts adheres strictly to data residency requirements, PII masking standards, and audit-log protocols. Define telemetry schema contracts with the OTel Instrumentation team to guarantee high upstream signal quality for downstream ML models. Author ML platform RFCs and contribute actively to observability data model standards across the broader engineering organization. Mentor junior ML and data engineers, and conduct rigorous design reviews for new mart schemas and model architectures. Required Skills: Kafka + Streaming (Flink/Spark) Lakehouse (Delta / Iceberg) ML (Anomaly detection + time-series) Observability (OTel, APM, Logs) MLOps (feature store, drift, retraining) SQL + Python (strong) Dynatrace Qualifications: 10+ years of experience
More at Prophecy Technologies