Source description
About the role
Our client is a NASDAQ listed company that employs over 24000 people globally. With an annual turnover of over 3bn Euro, they offer digital & software services to a wide range of clientele i.e., large scale enterprises to public sector companies in over 90 countries. They have development centres across the globe. In India, their presence is in Pune, Bangalore & Chandigarh. For the PUNE development centre, they are looking for a Data Engineer - Data Ingestion & Databricks Platform on CONTRACT TO HIRE. Data Engineer - Data Ingestion & Databricks Platform Exp Level : 6-9yrs Job location : Bangalore/ Chandigarh Mode : Job Description: We are looking for a skilled Data Engineer to join our Enterprise Data Management (EDM) team. You will design, build, and maintain data ingestion pipelines that move data from enterprise source systems Oracle, SQL Server, Salesforce, APIs) into AWS S3 and Databricks Unity Catalogue, following a Medallion architecture. Key Responsibilities Build and maintain batch and streaming data ingestion pipelines using PySpark on Databricks Manage CDC (Change Data Capture) pipelines using Striim application for Oracle and SQL Server sources Implement Delta Lake table operations - SCD Type 2, merge/upsert, Z-ordering, clustering Use Databricks Autoloader for S3 Bronze layer ingestion of Parquet, XML, and JSON files Develop post-load processing logic for sources like Salesforce, Google Analytics, XML. Manage pipeline metadata via a PostgreSQL config/catalog store (pipeline run logs, activity logs, error logs) Handle concurrent/multi-threaded pipeline execution using Python ThreadPoolExecutor Write unit tests and integration tests within Databricks notebooks Follow CI/CD workflows using Databricks Asset Bundles and Git-based branching (QA STG PROD) Troubleshoot pipeline failures with structured error handling, retry logic, and alerting Required Skills PySpark (DataFrames, SQL, Window functions, UDFs) Strong understanding of Medallion architecture (Bronze/Silver/Gold) Delta Lake (DeltaTable, merge, SCD2, Z-order, CLUSTER BY) Databricks (Notebooks, Jobs, Workflows, Widgets, Unity Catalog, Secrets, dbutils) AWS S3 (as data lake storage layer) Python (OOP, concurrent.futures, error handling, JSON/XML parsing) SQL (complex queries, PostgreSQL) XML/JSON data parsing Databricks Autoloader (structured streaming, multi-threaded) Structured Streaming (PySpark Streaming, checkpointing) Databricks Asset Bundles Git branching strategy, PRs, code reviews Unit & Integration Testing in Databricks Experience with Photon runtime and performance tuning on Databricks Familiarity with insurance domain data (policy, claims, Salesforce CRM) Experience migrating pipelines from Azure to AWS Databricks Knowledge of PostgreSQL as a metadata/catalog store Understanding of data governance (PII/HPII field handling, encryption) Experience with SFTP, FTPS, file-based ingestion patterns Familiarity with Databricks Unity Catalogue and foreign catalogues Candidates willing to work on a CONTRACT TO HIRE mode and available at short notice of no more than 20 days need only apply. Our client is a NASDAQ listed company that employs over 24000 people globally. With an annual turnover of over 3bn Euro, they offer digital & software services to a wide range of clientele i.e., large scale enterprises to public sector companies in over 90 countries. They have development centres across the globe. In India, their presence is in Pune, Bangalore & Chandigarh. For the PUNE development centre, they are looking for a Data Engineer - Data Ingestion & Databricks Platform on CONTRACT TO HIRE. Data Engineer - Data Ingestion & Databricks Platform Exp Level : 6-9yrs Job location : Bangalore/ Chandigarh Mode : Job Description: We are looking for a skilled Data Engineer to join our Enterprise Data Management (EDM) team. You will design, build, and maintain data ingestion pipelines that move data from enterprise source systems Oracle, SQL Server, Salesforce, APIs) into AWS S3 and Databricks Unity Catalogue, following a Medallion architecture. Key Responsibilities Build and maintain batch and streaming data ingestion pipelines using PySpark on Databricks Manage CDC (Change Data Capture) pipelines using Striim application for Oracle and SQL Server sources Implement Delta Lake table operations - SCD Type 2, merge/upsert, Z-ordering, clustering Use Databricks Autoloader for S3 Bronze layer ingestion of Parquet, XML, and JSON files Develop post-load processing logic for sources like Salesforce, Google Analytics, XML. Manage pipeline metadata via a PostgreSQL config/catalog store (pipeline run logs, activity logs, error logs) Handle concurrent/multi-threaded pipeline execution using Python ThreadPoolExecutor Write unit tests and integration tests within Databricks notebooks Follow CI/CD workflows using Databricks Asset Bundles and Git-based branching (QA STG PROD) Troubleshoot pipeline failures with structured error handling, retr
More at Eyeglobal