Padmi

Lead Data Engineer- GenAI/LLM (Mumbai)

MumbaiPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at ALTRAIZE

Opens the source posting on shine.com

Source description

About the role

View original

10+ years of hands-on data engineering experience, including technical leadership of data platform initiatives Strong, hands-on expertise with the Databricks Lakehouse Platform, including Delta Lake, Delta Live Tables, Databricks Workflows, and Unity Catalog Proven experience designing and implementing the medallion (bronze/silver/gold) architecture for data ingestion, curation, and consumption Strong expertise in distributed data processing using Apache Spark (PySpark and Spark SQL) or equivalent big-data frameworks Proven experience designing and building ETL/ELT pipelines for both batch and streaming data at scale Expert-level proficiency in Python and SQL for data transformation, validation, and pipeline development Solid experience building and managing data lakes and data warehouses, including dimensional and lakehouse data modeling Working knowledge of Large Language Models (LLMs) and GenAI concepts prompts, embeddings, vector databases, and Retrieval-Augmented Generation (RAG) and the data pipelines required to support them Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP) and its core data services Proven experience leading large-scale data migrations from on-premise (e.g. Hadoop) and cloud platforms to modern data architectures Strong understanding of distributed data processing, partitioning, and performance optimization techniques Experience implementing data security, governance, lineage, and access control (e.g. IAM, encryption, cataloging) Strong understanding of object-oriented programming, software design patterns, and CI/CD practices Familiarity with Agile/Scrum delivery methodologies and experience mentoring engineers Excellent problem-solving, analytical, and stakeholder-management skills Strong verbal and written communication skills Skills: pipelines,data,aws 10+ years of hands-on data engineering experience, including technical leadership of data platform initiatives Strong, hands-on expertise with the Databricks Lakehouse Platform, including Delta Lake, Delta Live Tables, Databricks Workflows, and Unity Catalog Proven experience designing and implementing the medallion (bronze/silver/gold) architecture for data ingestion, curation, and consumption Strong expertise in distributed data processing using Apache Spark (PySpark and Spark SQL) or equivalent big-data frameworks Proven experience designing and building ETL/ELT pipelines for both batch and streaming data at scale Expert-level proficiency in Python and SQL for data transformation, validation, and pipeline development Solid experience building and managing data lakes and data warehouses, including dimensional and lakehouse data modeling Working knowledge of Large Language Models (LLMs) and GenAI concepts prompts, embeddings, vector databases, and Retrieval-Augmented Generation (RAG) and the data pipelines required to support them Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP) and its core data services Proven experience leading large-scale data migrations from on-premise (e.g. Hadoop) and cloud platforms to modern data architectures Strong understanding of distributed data processing, partitioning, and performance optimization techniques Experience implementing data security, governance, lineage, and access control (e.g. IAM, encryption, cataloging) Strong understanding of object-oriented programming, software design patterns, and CI/CD practices Familiarity with Agile/Scrum delivery methodologies and experience mentoring engineers Excellent problem-solving, analytical, and stakeholder-management skills Strong verbal and written communication skills Skills: pipelines,data,aws

One address, no account. We’ll tell you when matching roles go live.

More at ALTRAIZE

Related open roles

View all roles