Padmi

Data Engg with Databricks - Technical Lead-Data Engg

IndiaPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Birlasoft

Opens the source posting on shine.com

Source description

About the role

View original

As an Azure Data Engineer at our company, you will play a crucial role in architecting, developing, and optimizing modern data platforms and analytical solutions on Azure. Your expertise in Azure Databricks, PySpark, ADF, and SQL will be essential for building enterprise-grade data pipelines, enabling ingestion from diverse sources, implementing complex transformations, and supporting scalable analytics initiatives. Key Responsibilities: - Data Solution Design & Architecture - Design end-to-end data engineering solutions using Azure Data Factory, Azure Databricks, PySpark, SQL, and other Azure-native services. - Architect and implement scalable and secure Modern Data Warehouse (MDW) and Lakehouse solutions leveraging Azure Data Lake Storage and Databricks. - Develop data models, integration patterns, and reusable frameworks aligned with best practices and enterprise architecture standards. - Participate in requirement discussions, solution blueprinting, and technical feasibility assessments. - Data Pipeline Development - Build and optimize robust, high-throughput ELT/ETL pipelines, enabling ingestion, transformation, and curation of structured, semi-structured, and unstructured data. - Integrate data from multiple on-premise and cloud-based systems, APIs, and third-party sources. - Implement complex transformations using PySpark, ensuring performance efficiency and code modularity. - Build orchestration workflows in ADF, including pipelines, triggers, linked services, integration runtimes, and parameterized datasets. - Databricks & PySpark Engineering - Develop scalable transformation scripts using PySpark on Databricks, applying advanced optimizations like caching, partitioning, and Delta Lake capabilities. - Implement Delta Lake featuresACID transactions, schema enforcement, schema evolution, and time travelacross the data lifecycle. - Perform performance tuning, handling bottlenecks related to cluster configuration, shuffle operations, joins, and parallelization. - Collaborate with platform teams to manage Databricks clusters, jobs, notebooks, and CI/CD integrations. - Data Governance, Quality & Security - Implement data quality checks, audit mechanisms, and validation frameworks to ensure data accuracy and consistency. - Ensure compliance with organizational standards for data security, encryption, access control, and data lifecycle management. - Create and maintain technical documentation, data flow diagrams, and operational support guides. - Collaboration & Stakeholder Management - Collaborate closely with BI, analytics, and business teams to understand data requirements and deliver reliable, production-ready solutions. - Work with architects, product owners, and cross-functional engineering teams to align technical delivery with business objectives. - Provide guidance and mentoring to junior engineers when required. As an Azure Data Engineer at our company, you will play a crucial role in architecting, developing, and optimizing modern data platforms and analytical solutions on Azure. Your expertise in Azure Databricks, PySpark, ADF, and SQL will be essential for building enterprise-grade data pipelines, enabling ingestion from diverse sources, implementing complex transformations, and supporting scalable analytics initiatives. Key Responsibilities: - Data Solution Design & Architecture - Design end-to-end data engineering solutions using Azure Data Factory, Azure Databricks, PySpark, SQL, and other Azure-native services. - Architect and implement scalable and secure Modern Data Warehouse (MDW) and Lakehouse solutions leveraging Azure Data Lake Storage and Databricks. - Develop data models, integration patterns, and reusable frameworks aligned with best practices and enterprise architecture standards. - Participate in requirement discussions, solution blueprinting, and technical feasibility assessments. - Data Pipeline Development - Build and optimize robust, high-throughput ELT/ETL pipelines, enabling ingestion, transformation, and curation of structured, semi-structured, and unstructured data. - Integrate data from multiple on-premise and cloud-based systems, APIs, and third-party sources. - Implement complex transformations using PySpark, ensuring performance efficiency and code modularity. - Build orchestration workflows in ADF, including pipelines, triggers, linked services, integration runtimes, and parameterized datasets. - Databricks & PySpark Engineering - Develop scalable transformation scripts using PySpark on Databricks, applying advanced optimizations like caching, partitioning, and Delta Lake capabilities. - Implement Delta Lake featuresACID transactions, schema enforcement, schema evolution, and time travelacross the data lifecycle. - Perform performance tuning, handling bottlenecks related to cluster configuration, shuffle operations, joins, and parallelization. - Collabor

One address, no account. We’ll tell you when matching roles go live.

More at Birlasoft

Related open roles

View all roles