Padmi
NTT DATA logo
NTT DATA

SAP implementation and managed services · ServiceNow consulting and FSO

Data Engineer Sr (Databricks)

BangalorePosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at NTT DATA

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Databricks Engineer, your role involves developing Databricks notebooks, jobs, and workflows to replicate and enhance DB2/Guidewire-based pipelines and transformations. You will be responsible for implementing Delta Lake tables and patterns (bronze/silver/gold, ACID, time travel, schema evolution) for migrated data. Additionally, you will integrate Databricks with AWS/S3 or Azure ADLS, ADF/Synapse, Key Vault, and Snowflake as required. Your responsibilities also include optimizing Databricks clusters, jobs, and queries for performance and cost, implementing incremental loads, CDC patterns, and batch schedules for large datasets, and collaborating with Snowflake and dbt teams to ensure consistent data models and data contracts. Qualifications Required: - 8+ years of experience in Databricks with a focus on Legacy Demystification & Ingestion, specifically in DB2/400 & Guidewire environments - Experience in extracting DB2/AS400 data using Change Data Capture (CDC) or scheduled batch extractions into cloud storage - Proficiency in handling Guidewire data by integrating with Guidewire Cloud Data Access (CDA) or InsuranceSuite to replicate complex P&C insurance schemas - Expertise in Delta Lake optimization, schema evolution, upserts, and SCD Type 2 changes using Databricks and Apache Spark - Strong background in refactoring legacy procedural code into scalable distributed patterns using PySpark, Spark SQL, and Scala - Knowledge of data governance, lineage tracing, and data quality automation - Experience in managing Databricks serverless resources, building event-driven pipelines, and ingesting flat files and streaming records into Delta tables In this role, you will play a crucial part in governing vast amounts of incoming and generated data across the enterprise by implementing strict data governance, lineage tracing, and table-level security using Unity Catalog. You will also be responsible for automating data validation frameworks to ensure a seamless transition from legacy to modern systems without data loss or corruption. Additionally, you will be involved in managing Databricks serverless resources, ensuring optimal cluster sizing, and reducing compute costs, as well as building event-driven pipelines using features like Databricks Auto Loader to ingest flat files and streaming records directly into Delta tables. As a Senior Databricks Engineer, your role involves developing Databricks notebooks, jobs, and workflows to replicate and enhance DB2/Guidewire-based pipelines and transformations. You will be responsible for implementing Delta Lake tables and patterns (bronze/silver/gold, ACID, time travel, schema evolution) for migrated data. Additionally, you will integrate Databricks with AWS/S3 or Azure ADLS, ADF/Synapse, Key Vault, and Snowflake as required. Your responsibilities also include optimizing Databricks clusters, jobs, and queries for performance and cost, implementing incremental loads, CDC patterns, and batch schedules for large datasets, and collaborating with Snowflake and dbt teams to ensure consistent data models and data contracts. Qualifications Required: - 8+ years of experience in Databricks with a focus on Legacy Demystification & Ingestion, specifically in DB2/400 & Guidewire environments - Experience in extracting DB2/AS400 data using Change Data Capture (CDC) or scheduled batch extractions into cloud storage - Proficiency in handling Guidewire data by integrating with Guidewire Cloud Data Access (CDA) or InsuranceSuite to replicate complex P&C insurance schemas - Expertise in Delta Lake optimization, schema evolution, upserts, and SCD Type 2 changes using Databricks and Apache Spark - Strong background in refactoring legacy procedural code into scalable distributed patterns using PySpark, Spark SQL, and Scala - Knowledge of data governance, lineage tracing, and data quality automation - Experience in managing Databricks serverless resources, building event-driven pipelines, and ingesting flat files and streaming records into Delta tables In this role, you will play a crucial part in governing vast amounts of incoming and generated data across the enterprise by implementing strict data governance, lineage tracing, and table-level security using Unity Catalog. You will also be responsible for automating data validation frameworks to ensure a seamless transition from legacy to modern systems without data loss or corruption. Additionally, you will be involved in managing Databricks serverless resources, ensuring optimal cluster sizing, and reducing compute costs, as well as building event-driven pipelines using features like Databricks Auto Loader to ingest flat files and streaming records directly into Delta tables.

One address, no account. We’ll tell you when matching roles go live.

More at NTT DATA

Related open roles

View all roles