Padmi

Data Engineer Ingestion & Pipelines

MumbaiPosted 2 months ago
Data Science And StatisticsMid-levelFull Time; Regular
Apply at Programming.Com

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Engineer specializing in Ingestion & Pipelines at Reliance Enterprise Intelligence Ltd (REIL), your primary responsibility is to establish and maintain the data foundation of the enterprise AI platform. The success of the platform in making critical business decisions relies heavily on the accuracy, completeness, and timeliness of the data you manage. Key Responsibilities: - Data Discovery & Audit - Collaborate with internal IT and operations teams to conduct comprehensive data audits, mapping all necessary source systems and evaluating existing platform data. - Maintain detailed documentation of data sources, schemas, update frequencies, and known data quality constraints. - Proactively identify and escalate data gaps and quality risks to the Solutions Architect. - Pipeline Design & Build - Design and implement high-throughput ingestion pipelines from diverse sources such as ERPs, government portals, supplier networks, and banking feeds into the data platform. - Implement robust batch processing and low-latency real-time/streaming ingestion patterns. - Ensure fault tolerance in pipelines to handle system failures gracefully without data loss or duplication. - Data Quality & Reliability - Establish proactive pipeline monitoring and alerting frameworks to detect and address failures before affecting model training or live inference. - Implement inline data quality checks at ingestion points, including schema validation, completeness verification, and anomaly detection. - Maintain clear data lineage for transparency on data origin and freshness. - Collaboration & Handoff - Collaborate closely with internal IT and automation teams to leverage knowledge of legacy source systems. - Deliver clean, optimized datasets to ML and LLM Engineers for model training and knowledge base construction. - Support the MLOps Engineer in ensuring stable and performant production pipelines post-deployment. Qualifications & Experience: - Education: B.E./B.Tech/M.Tech in Computer Science, Information Technology, or a related technical field. - Required Experience: - 5+ years of data engineering experience, with at least 2 years focused on building enterprise-scale production pipelines. - Demonstrated expertise in extracting data from complex enterprise systems. - Core Technical Skills: - Proficiency in Python, PySpark, and advanced SQL. - Hands-on experience with Databricks and Delta Lake. - Familiarity with Apache Airflow, Kafka, Spark Structured Streaming, ERP integration, API management, data quality frameworks, and observability tools. - Preferred Qualifications: - Deep understanding of SAP data models. - Experience with government API ecosystems. - Exposure to data cataloging and governance tools. - Background in fintech, compliance, or tax technology environments. In addition to the technical qualifications, you will be responsible for ensuring the stability and efficiency of production pipelines and collaborating with various teams to optimize data processes. Your role will be pivotal in driving the success of the enterprise AI platform at REIL. As a Data Engineer specializing in Ingestion & Pipelines at Reliance Enterprise Intelligence Ltd (REIL), your primary responsibility is to establish and maintain the data foundation of the enterprise AI platform. The success of the platform in making critical business decisions relies heavily on the accuracy, completeness, and timeliness of the data you manage. Key Responsibilities: - Data Discovery & Audit - Collaborate with internal IT and operations teams to conduct comprehensive data audits, mapping all necessary source systems and evaluating existing platform data. - Maintain detailed documentation of data sources, schemas, update frequencies, and known data quality constraints. - Proactively identify and escalate data gaps and quality risks to the Solutions Architect. - Pipeline Design & Build - Design and implement high-throughput ingestion pipelines from diverse sources such as ERPs, government portals, supplier networks, and banking feeds into the data platform. - Implement robust batch processing and low-latency real-time/streaming ingestion patterns. - Ensure fault tolerance in pipelines to handle system failures gracefully without data loss or duplication. - Data Quality & Reliability - Establish proactive pipeline monitoring and alerting frameworks to detect and address failures before affecting model training or live inference. - Implement inline data quality checks at ingestion points, including schema validation, completeness verification, and anomaly detection. - Maintain clear data lineage for transparency on data origin and freshness. - Collaboration & Handoff - Collaborate closely with internal IT and automation teams to leverage knowledge of legacy source systems. - Deliver clean, optimized datasets to ML and LLM Engi

One address, no account. We’ll tell you when matching roles go live.

More at Programming.Com

Related open roles

View all roles