Source description
About the role
Purpose/Objective The Databricks Engineer will build, test, optimise and support Databricks-based data pipelines and Lakehouse components. The role requires strong hands-on Python coding capability, PySpark development skills and practical experience in production-grade data engineering. Key Responsibilities of Role Key Responsibilities - Develop data pipelines using Python, PySpark, SQL and Databricks notebooks/jobs. - Build ETL/ELT workflows for ingestion, transformation, validation and publishing of data. - Implement Delta Lake tables, partitioning strategies and optimised query patterns. - Perform unit testing, data reconciliation, exception handling and production issue support. - Work with senior engineers and architects to implement standardised Lakehouse patterns. - Participate in code reviews, sprint delivery, release activities and documentation. - Support CI/CD integration, Git-based development and deployment practices. - Monitor pipeline failures, troubleshoot performance issues and support operational runbooks. Mandatory Skills - Python coding is mandatory: candidate must clear a hands-on Python coding round. - Strong Python programming knowledge including functions, classes, exception handling and file processing. - Hands-on PySpark development experience. - Databricks notebooks, jobs/workflows and cluster usage experience. - Good SQL skills and understanding of joins, aggregations, window functions and query optimisation. - Experience with ETL/ELT, data transformation and data validation logic. - Basic understanding of Git and CI/CD practices. Python Coding Assessment Scope - Data transformation using lists, dictionaries, files or dataframes. - Writing reusable functions and clean modular code. - Handling missing, duplicate or invalid data records. - Basic object-oriented programming concepts. - Exception handling and logging approach. - Simple API/file ingestion or parsing problem. - Optional PySpark coding problem for O3 candidates. Preferred Skills - Delta Lake, Unity Catalog, Delta Live Tables or structured streaming exposure. - Cloud platform experience on Azure or Google Cloud Platform. - Data quality, metadata, lineage and governance awareness. - Terraform, Databricks Asset Bundles or deployment automation exposure. - Databricks Associate certification will be an advantage. Selection Criteria - Python coding assessment: mandatory and eliminatory. - Technical interview covering Databricks, PySpark, SQL and production data engineering. - Scenario discussion on pipeline failure, performance tuning and data quality issue handling. - Final discussion on communication, ownership, learning ability and team fit. Technical Competencies Data Engineering & ETL Development,Databricks Lakehouse & Delta Lake Architecture,Data Quality, Testing & Production Support,CI/CD, DevOps & Operational Excellence Qualifications and Experience 2-3 years overall experience, with hands-on Python, PySpark and Databricks delivery exposure
More at Adani