Source description
About the role
Role & responsibilities Candidate should have 4-7 years of experience in data engineering, DataOps, or platform engineering roles Strong proficiency in Python and SQL for data pipeline development Experience with time-series databases (OSI PI, Honeywell PHD, InfluxDB) Familiarity with CDF platform Experience with industrial data sources (OPC-UA, MQTT, Modbus, historians Experience with Azure datalake and data warehouse, IIoT services Experience in connecting with SAP S/4Hana, ECC PM, MM modules and ingesting batch data through delivery pipelines Experience with data pipeline orchestration tools (Apache Airflow, Dagster, Prefect, or Azure Data Factory) Proficiency with stream processing frameworks (Kafka, Spark Streaming, Flink, or Azure Event Hubs) Experience with data warehousing and data lake solutions (Snowflake, Databricks, Azure Synapse) Strong knowledge of Docker and containerization Familiarity with infrastructure as code (Terraform, ARM, Bicep) and CI/CD pipelines Essential Duties and Key Competencies : Design, build, and maintain scalable data pipelines for ingesting industrial time-series data from sensors, historians, and IoT devices Develop and operate ETL/ELT processes for batch and streaming data from diverse sources (SAP, CMMS, inspection reports, documents) Build and optimize time-series database schemas for high-velocity industrial data (millions of data points per minute) Implement data validation, data quality checks, and monitoring for all data pipelines Deploy and manage data infrastructure on cloud platforms (Azure preferred, AWS) Orchestrate complex data workflows using Airflow, custom connectors or Azure Data Factory Collaborate with data scientists and ML engineers to provision data for model training and inference Implement data partitioning, sharding, and retention policies for terabyte-scale datasets Build and maintain APIs for data serving to downstream applications and AI models Ensure data security, encryption, and access controls across all data stores Monitor pipeline performance, troubleshoot failures, and optimize for latency and cost Document data lineage, data dictionaries, and pipeline architectures Additional Skills (Optional) Experience with vector databases (Pinecone, Weaviate, Milvus) for RAG applications Familiarity with feature stores (Feast, Tecton, Databricks Feature Store) Experience with data version control (DVC, LakeFS) Knowledge of MLOps practices and model data pipelines Experience with industrial protocols (OPC-UA, MQTT, Modbus) and historian systems Understanding of data governance and compliance (GDPR, ISO 27001) Experience with real-time anomaly detection pipelines Preferred candidate profile
More at Tridiagonal Ai