Padmi
Innodata logo
Innodata

generative AI training data · data annotation

GCP Data Engineer

IndiaPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Innodata

Opens the source posting on shine.com

Source description

About the role

View original

You will be responsible for designing and implementing data-driven solutions on GCP, including BigQuery, Cloud Storage, Dataflow, Pub/Sub, and Looker/BI. Your key responsibilities will include: - Building ETL scripts using SQL and Python to extract, clean, and transform structured and unstructured data from ERP, procurement, logistics, and facility management systems. - Developing and optimizing data pipelines for ingestion, transformation, and loading into enterprise data lakes and warehouses. - Building and extending end-to-end data and BI solutions, covering extraction, storage, transformation, and visualization layers. - Partnering with supply chain, real estate, and AI/ML teams to provide pipelines for AI solutions such as RAG ingestion, Copilot integration, and multi-agent workflows. - Ensuring data governance, lineage, and compliance across supply chain datasets. - Continuously optimizing query performance, ETL processes, and pipeline reliability. The skills required for this role include: - Strong hands-on expertise with GCP services like BigQuery, Dataflow, Pub/Sub, Cloud Storage, Looker/BI, or similar. - Advanced proficiency in SQL for complex queries and optimization, as well as Python for data engineering, scripting, and APIs. - Experience in building ETL/ELT pipelines operating on structured and unstructured data sources. - Knowledge of enterprise data warehouse and data lake architectures. If there are any additional details about the company in the job description, please provide them so that I can include them in the job description. You will be responsible for designing and implementing data-driven solutions on GCP, including BigQuery, Cloud Storage, Dataflow, Pub/Sub, and Looker/BI. Your key responsibilities will include: - Building ETL scripts using SQL and Python to extract, clean, and transform structured and unstructured data from ERP, procurement, logistics, and facility management systems. - Developing and optimizing data pipelines for ingestion, transformation, and loading into enterprise data lakes and warehouses. - Building and extending end-to-end data and BI solutions, covering extraction, storage, transformation, and visualization layers. - Partnering with supply chain, real estate, and AI/ML teams to provide pipelines for AI solutions such as RAG ingestion, Copilot integration, and multi-agent workflows. - Ensuring data governance, lineage, and compliance across supply chain datasets. - Continuously optimizing query performance, ETL processes, and pipeline reliability. The skills required for this role include: - Strong hands-on expertise with GCP services like BigQuery, Dataflow, Pub/Sub, Cloud Storage, Looker/BI, or similar. - Advanced proficiency in SQL for complex queries and optimization, as well as Python for data engineering, scripting, and APIs. - Experience in building ETL/ELT pipelines operating on structured and unstructured data sources. - Knowledge of enterprise data warehouse and data lake architectures. If there are any additional details about the company in the job description, please provide them so that I can include them in the job description.

One address, no account. We’ll tell you when matching roles go live.

More at Innodata

Related open roles

View all roles