Padmi

Data Scientist / Data Engineer

MumbaiPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Celebal Technologies

Opens the source posting on shine.com

Source description

About the role

View original

As a highly skilled Azure Data Engineer, your role will involve designing and implementing streaming data pipelines that integrate Kafka with Databricks using Structured Streaming. You will also be responsible for architecting and maintaining Medallion Architecture with well-defined Bronze, Silver, and Gold layers. Additionally, you will implement efficient data ingestion using Databricks Autoloader for high-throughput data loads and work with large volumes of structured and unstructured data to ensure high availability and performance. Performance tuning techniques such as partitioning, caching, and cluster resource optimization will be applied, and collaboration with cross-functional teams to build robust data solutions will be essential. Establishing best practices for code versioning, deployment automation, and data governance will also be part of your responsibilities. Key Responsibilities: - Design and implement streaming data pipelines integrating Kafka with Databricks using Structured Streaming - Architect and maintain Medallion Architecture with well-defined Bronze, Silver, and Gold layers - Implement efficient ingestion using Databricks Autoloader for high-throughput data loads - Work with large volumes of structured and unstructured data, ensuring high availability and performance - Apply performance tuning techniques such as partitioning, caching, and cluster resource optimization - Collaborate with cross-functional teams (data scientists, analysts, business users) to build robust data solutions - Establish best practices for code versioning, deployment automation, and data governance Required Technical Skills: - Strong expertise in Azure Databricks and Spark Structured Streaming - 3-8 years of experience in Data Engineering - Processing modes (append, update, complete) - Output modes (append, complete, update) - Checkpointing and state management - Experience with Kafka integration for real-time data pipelines - Deep understanding of Medallion Architecture - Proficiency with Databricks Autoloader and schema evolution - Deep understanding of Unity Catalog and Foreign catalog - Strong knowledge of Spark SQL, Delta Lake, and DataFrames - Expertise in performance tuning (query optimization, cluster configuration, caching strategies) - Must have Data management strategies - Excellent with Governance and Access management - Strong with Data modelling, Data warehousing concepts, Databricks as a platform - Solid understanding of Window functions As a highly skilled Azure Data Engineer, your role will involve designing and implementing streaming data pipelines that integrate Kafka with Databricks using Structured Streaming. You will also be responsible for architecting and maintaining Medallion Architecture with well-defined Bronze, Silver, and Gold layers. Additionally, you will implement efficient data ingestion using Databricks Autoloader for high-throughput data loads and work with large volumes of structured and unstructured data to ensure high availability and performance. Performance tuning techniques such as partitioning, caching, and cluster resource optimization will be applied, and collaboration with cross-functional teams to build robust data solutions will be essential. Establishing best practices for code versioning, deployment automation, and data governance will also be part of your responsibilities. Key Responsibilities: - Design and implement streaming data pipelines integrating Kafka with Databricks using Structured Streaming - Architect and maintain Medallion Architecture with well-defined Bronze, Silver, and Gold layers - Implement efficient ingestion using Databricks Autoloader for high-throughput data loads - Work with large volumes of structured and unstructured data, ensuring high availability and performance - Apply performance tuning techniques such as partitioning, caching, and cluster resource optimization - Collaborate with cross-functional teams (data scientists, analysts, business users) to build robust data solutions - Establish best practices for code versioning, deployment automation, and data governance Required Technical Skills: - Strong expertise in Azure Databricks and Spark Structured Streaming - 3-8 years of experience in Data Engineering - Processing modes (append, update, complete) - Output modes (append, complete, update) - Checkpointing and state management - Experience with Kafka integration for real-time data pipelines - Deep understanding of Medallion Architecture - Proficiency with Databricks Autoloader and schema evolution - Deep understanding of Unity Catalog and Foreign catalog - Strong knowledge of Spark SQL, Delta Lake, and DataFrames - Expertise in performance tuning (query optimization, cluster configuration, caching strategies) - Must have Data management strategies - Excellent with Governance and Access management - Strong with Data modelling, Data warehousing concepts, Databricks as a platform - Solid understanding of Window functions

One address, no account. We’ll tell you when matching roles go live.

More at Celebal Technologies

Related open roles

View all roles