Source description
About the role
Job Summary We are seeking an experienced Data Architect to lead the modernization of our data platforms. You will own the architectural strategy for migrating complex data workflows from DataIKU to Azure Databricks, ensuring scalable, high-performance, and cost-efficient pipeline design. Note: Immediate joiners are preferred Key Responsibilities Lead the end-to-end migration of legacy data workflows from DataIKU to Azure Databricks. Evaluate and map current DataIKU data streams and infrastructure. Refactor and enhance data processing through PySpark-driven Databricks solutions. Design robust, scalable ETL/ELT architectures using ADF and Databricks. Optimize Spark jobs, query performance, and storage strategies (Delta Lake/ADLS). Architect and deploy robust data modeling solutions utilizing Delta Lake and ADLS. Refine storage efficiency through strategic implementation of: Data partitioning techniques Advanced file formats including Parquet and Delta Performance Engineering: Maximize Spark job throughput while minimizing operational expenditure. Calibrate complex queries and pipelines utilizing advanced caching and join techniques. Implement automation, CI/CD pipelines, and best practices for data validation and consistency. Document architectural standards and mentor teams through the transition and KT process. Validation & Testing: Ensure data consistency between DataIKU and Databricks outputs. Develop and execute reconciliation and validation scripts Required Experience Total Experience: 12 to 16 yrs in Data Engineering. Domain Expertise: Proven experience leading data platform migration projects (lift-and-shift/re-platforming). Specialization: 3+ years of hands-on experience in the Databricks/Spark ecosystem. Tech Stack Cloud & Processing: Azure Databricks (Mandatory), Azure Data Factory (ADF), Azure Data Lake Storage (ADLS Gen2). Languages & Tools: Python (Advanced), PySpark, SQL (Advanced), Delta Lake architecture. DevOps: Azure DevOps / GitHub Actions (CI/CD), Git. Legacy Tooling: Experience with DataIKU (DSS). Key Deliverables Execution of end-to-end pipeline migrations from DataIKU to Databricks. Deployment of scalable ADF orchestration workflows. High-throughput PySpark jobs refined for performance. Comprehensive data reconciliation and validation audit reports. Architectural standards, runbooks, and knowledge transfer documentation.
More at Infogain