Source description
About the role
As a Data Engineer, you will be responsible for designing and implementing scalable batch and streaming data pipelines using Azure Databricks, Delta Live Tables, and Apache Spark. Your key responsibilities will include: - Building and maintaining ETL/ELT workflows orchestrated through Azure Data Factory and Databricks Workflows. - Developing reusable and modular pipeline components following software engineering best practices. - Architecting and managing Lakehouse solutions using Delta Lake across Bronze, Silver, and Gold layers. - Designing and enforcing data models, schemas, and governance policies using Unity Catalog. - Optimizing storage, partitioning, and query performance for large-scale datasets on ADLS Gen2. - Managing Databricks clusters, compute policies, and job scheduling. - Implementing Infrastructure as Code (IaC) using Terraform or ARM templates. - Integrating Databricks with Azure services including Synapse, Event Hubs, Key Vault, and Azure DevOps. Qualifications required for this role include: - Strong proficiency in PySpark, Python, and SQL for Big Data processing. - Hands-on experience with Delta Lake, Delta Live Tables (DLT), and Medallion Architecture. - Strong experience with Azure Data Services including Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Data Factory (ADF), Azure Synapse Analytics, and Azure Event Hubs. - Experience with Databricks Unity Catalog for data governance and access control. - Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions for Databricks deployments. - Strong understanding of distributed computing concepts, Spark optimization, partitioning, and performance tuning. - Experience with streaming data processing using Structured Streaming, Kafka, or Event Hubs. As a Data Engineer, you will be responsible for designing and implementing scalable batch and streaming data pipelines using Azure Databricks, Delta Live Tables, and Apache Spark. Your key responsibilities will include: - Building and maintaining ETL/ELT workflows orchestrated through Azure Data Factory and Databricks Workflows. - Developing reusable and modular pipeline components following software engineering best practices. - Architecting and managing Lakehouse solutions using Delta Lake across Bronze, Silver, and Gold layers. - Designing and enforcing data models, schemas, and governance policies using Unity Catalog. - Optimizing storage, partitioning, and query performance for large-scale datasets on ADLS Gen2. - Managing Databricks clusters, compute policies, and job scheduling. - Implementing Infrastructure as Code (IaC) using Terraform or ARM templates. - Integrating Databricks with Azure services including Synapse, Event Hubs, Key Vault, and Azure DevOps. Qualifications required for this role include: - Strong proficiency in PySpark, Python, and SQL for Big Data processing. - Hands-on experience with Delta Lake, Delta Live Tables (DLT), and Medallion Architecture. - Strong experience with Azure Data Services including Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Data Factory (ADF), Azure Synapse Analytics, and Azure Event Hubs. - Experience with Databricks Unity Catalog for data governance and access control. - Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions for Databricks deployments. - Strong understanding of distributed computing concepts, Spark optimization, partitioning, and performance tuning. - Experience with streaming data processing using Structured Streaming, Kafka, or Event Hubs.
More at InFynd