Padmi

Data Engineer Azure Databricks (Mumbai)

MumbaiPosted 1 month ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at Infinivisglobal

Opens the source posting on shine.com

Source description

About the role

View original

About the Role We are seeking a highly skilled Data Engineer with strong expertise in Azure Databricks, SQL, PySpark, and Data Modeling. The ideal candidate will have experience designing and implementing scalable data pipelines, optimizing data workflows, and building modern data platforms on the Azure ecosystem. *Key Responsibilities - Design, develop, and maintain ETL/ELT pipelines using Azure Databricks & PySpark. - Build and manage Delta Lakehouse solutions including Bronze, Silver, Gold layers. - Collaborate with data architects, analysts, and business stakeholders to design data models (Star/Snowflake schemas, Fact & Dimension tables). - Optimize Databricks clusters, jobs, and queries for performance and cost-efficiency. - Implement CI/CD pipelines for Databricks notebooks and data workflows. - Manage schema evolution, data governance, and quality checks across pipelines. - Work with Azure Data Lake Storage (ADLS), Azure Synapse Analytics, and SQL Databases for end-to-end data solutions. - Implement data partitioning, caching, and broadcast joins to optimize PySpark jobs. - Ensure best practices in data security, compliance, and access management. - Troubleshoot and optimize slow SQL queries, indexes, and data warehouse performance. - Support business reporting and analytics needs by designing and maintaining scalable data models. Requirements *Required Skills - Azure Databricks: Notebooks, Delta Tables, Auto-scaling, Job Orchestration. - SQL: Joins, Window Functions, Indexing, Query Optimization, CTEs, SCD handling. - PySpark: RDD, Data Frame API, Lazy evaluation, Transformations, Optimizations. - Data Modeling: OLTP vs OLAP, Star & Snowflake Schema, Fact/Dimension Tables, Normalization/Denormalization. - Azure Ecosystem: ADLS, Synapse Analytics, Azure Data Factory (ADF) is a plus. - Solid understanding of ETL best practices, data quality frameworks, and large-scale distributed data processing. Nice-to-Have Skills - Experience with DataBricks CI/CD pi About the Role We are seeking a highly skilled Data Engineer with strong expertise in Azure Databricks, SQL, PySpark, and Data Modeling. The ideal candidate will have experience designing and implementing scalable data pipelines, optimizing data workflows, and building modern data platforms on the Azure ecosystem. *Key Responsibilities - Design, develop, and maintain ETL/ELT pipelines using Azure Databricks & PySpark. - Build and manage Delta Lakehouse solutions including Bronze, Silver, Gold layers. - Collaborate with data architects, analysts, and business stakeholders to design data models (Star/Snowflake schemas, Fact & Dimension tables). - Optimize Databricks clusters, jobs, and queries for performance and cost-efficiency. - Implement CI/CD pipelines for Databricks notebooks and data workflows. - Manage schema evolution, data governance, and quality checks across pipelines. - Work with Azure Data Lake Storage (ADLS), Azure Synapse Analytics, and SQL Databases for end-to-end data solutions. - Implement data partitioning, caching, and broadcast joins to optimize PySpark jobs. - Ensure best practices in data security, compliance, and access management. - Troubleshoot and optimize slow SQL queries, indexes, and data warehouse performance. - Support business reporting and analytics needs by designing and maintaining scalable data models. Requirements *Required Skills - Azure Databricks: Notebooks, Delta Tables, Auto-scaling, Job Orchestration. - SQL: Joins, Window Functions, Indexing, Query Optimization, CTEs, SCD handling. - PySpark: RDD, Data Frame API, Lazy evaluation, Transformations, Optimizations. - Data Modeling: OLTP vs OLAP, Star & Snowflake Schema, Fact/Dimension Tables, Normalization/Denormalization. - Azure Ecosystem: ADLS, Synapse Analytics, Azure Data Factory (ADF) is a plus. - Solid understanding of ETL best practices, data quality frameworks, and large-scale distributed data processing. Nice-to-Have Skills - Experience with DataBricks CI/CD pi

One address, no account. We’ll tell you when matching roles go live.

More at Infinivisglobal

Related open roles

View all roles