Source description
About the role
As a Data Engineer at a fast-scaling technology consultancy in the Enterprise Data & AI Engineering space, you will be responsible for designing and deploying production-grade data pipelines using Azure Databricks, PySpark, and Delta Lake for batch and streaming workloads. Your key responsibilities will include: - Designing and deploying production-grade data pipelines using Azure Databricks, PySpark, and Delta Lake for batch and streaming workloads. - Optimizing Databricks clusters for cost, performance, and concurrency. - Implementing CI/CD workflows via Azure DevOps or GitHub Actions. - Integrating Databricks with Azure Data Factory, Synapse, ADLS Gen2, and Power BI to enable end-to-end analytics and ML lifecycle. - Enforcing data governance and lineage using Unity Catalog, managing secrets via Key Vault, and implementing RBAC for secure multi-tenant workspaces. - Collaborating with Data Scientists to productionize ML modelsfeature engineering, training, and batch scoring pipelines. - Documenting architecture patterns, performance benchmarking, and operational runbooks for client handover and internal knowledge sharing. Qualifications Required: - Proficiency in Azure Databricks, PySpark, Delta Lake, Azure Data Factory, ADLS Gen2, Unity Catalog, CI/CD (Azure DevOps / GitHub Actions), and SQL. In this role, you will have the opportunity to work with Fortune 500 clients across global industries and solve real-world data challenges. You will experience an on-site role with a collaborative, high-ownership culture that values fast execution and provides access to certified Azure training, cloud credits, and mentorship from senior engineers. As a Data Engineer at a fast-scaling technology consultancy in the Enterprise Data & AI Engineering space, you will be responsible for designing and deploying production-grade data pipelines using Azure Databricks, PySpark, and Delta Lake for batch and streaming workloads. Your key responsibilities will include: - Designing and deploying production-grade data pipelines using Azure Databricks, PySpark, and Delta Lake for batch and streaming workloads. - Optimizing Databricks clusters for cost, performance, and concurrency. - Implementing CI/CD workflows via Azure DevOps or GitHub Actions. - Integrating Databricks with Azure Data Factory, Synapse, ADLS Gen2, and Power BI to enable end-to-end analytics and ML lifecycle. - Enforcing data governance and lineage using Unity Catalog, managing secrets via Key Vault, and implementing RBAC for secure multi-tenant workspaces. - Collaborating with Data Scientists to productionize ML modelsfeature engineering, training, and batch scoring pipelines. - Documenting architecture patterns, performance benchmarking, and operational runbooks for client handover and internal knowledge sharing. Qualifications Required: - Proficiency in Azure Databricks, PySpark, Delta Lake, Azure Data Factory, ADLS Gen2, Unity Catalog, CI/CD (Azure DevOps / GitHub Actions), and SQL. In this role, you will have the opportunity to work with Fortune 500 clients across global industries and solve real-world data challenges. You will experience an on-site role with a collaborative, high-ownership culture that values fast execution and provides access to certified Azure training, cloud credits, and mentorship from senior engineers.
More at Zorba AI