Padmi

Data Engineer (Big Data)

IndiaPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at InFynd

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Engineer, you will be responsible for designing and implementing scalable batch and streaming data pipelines using Azure Databricks, Delta Live Tables, and Apache Spark. Your key responsibilities will include: - Building and maintaining ETL/ELT workflows orchestrated through Azure Data Factory and Databricks Workflows. - Developing reusable and modular pipeline components following software engineering best practices. - Architecting and managing Lakehouse solutions using Delta Lake across Bronze, Silver, and Gold layers. - Designing and enforcing data models, schemas, and governance policies using Unity Catalog. - Optimizing storage, partitioning, and query performance for large-scale datasets on ADLS Gen2. - Managing Databricks clusters, compute policies, and job scheduling. - Implementing Infrastructure as Code (IaC) using Terraform or ARM templates. - Integrating Databricks with Azure services including Synapse, Event Hubs, Key Vault, and Azure DevOps. Qualifications required for this role include: - Strong proficiency in PySpark, Python, and SQL for Big Data processing. - Hands-on experience with Delta Lake, Delta Live Tables (DLT), and Medallion Architecture. - Strong experience with Azure Data Services including Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Data Factory (ADF), Azure Synapse Analytics, and Azure Event Hubs. - Experience with Databricks Unity Catalog for data governance and access control. - Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions for Databricks deployments. - Strong understanding of distributed computing concepts, Spark optimization, partitioning, and performance tuning. - Experience with streaming data processing using Structured Streaming, Kafka, or Event Hubs. As a Data Engineer, you will be responsible for designing and implementing scalable batch and streaming data pipelines using Azure Databricks, Delta Live Tables, and Apache Spark. Your key responsibilities will include: - Building and maintaining ETL/ELT workflows orchestrated through Azure Data Factory and Databricks Workflows. - Developing reusable and modular pipeline components following software engineering best practices. - Architecting and managing Lakehouse solutions using Delta Lake across Bronze, Silver, and Gold layers. - Designing and enforcing data models, schemas, and governance policies using Unity Catalog. - Optimizing storage, partitioning, and query performance for large-scale datasets on ADLS Gen2. - Managing Databricks clusters, compute policies, and job scheduling. - Implementing Infrastructure as Code (IaC) using Terraform or ARM templates. - Integrating Databricks with Azure services including Synapse, Event Hubs, Key Vault, and Azure DevOps. Qualifications required for this role include: - Strong proficiency in PySpark, Python, and SQL for Big Data processing. - Hands-on experience with Delta Lake, Delta Live Tables (DLT), and Medallion Architecture. - Strong experience with Azure Data Services including Azure Data Lake Storage Gen2 (ADLS Gen2), Azure Data Factory (ADF), Azure Synapse Analytics, and Azure Event Hubs. - Experience with Databricks Unity Catalog for data governance and access control. - Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions for Databricks deployments. - Strong understanding of distributed computing concepts, Spark optimization, partitioning, and performance tuning. - Experience with streaming data processing using Structured Streaming, Kafka, or Event Hubs.

One address, no account. We’ll tell you when matching roles go live.

More at InFynd

Related open roles

View all roles