Source description
About the role
Responsibilities: Design, develop, and maintain scalable and efficient data pipelines using PySpark , Scala Spark , Databricks , Python , and SQL . Write optimized, reusable, and high-quality code for data processing and transformation. Optimize SQL queries for high-performance data extraction, manipulation, and analysis. Demonstrate strong expertise in Databricks , including workflow management , job orchestration , and data exploration . Collaborate with cross-functional teams to gather and understand business and data requirements. Implement best practices for ETL , data pipeline optimization , and query performance tuning . Develop and maintain comprehensive documentation for all data pipelines, workflows, and related processes. Troubleshoot, debug, and resolve data pipeline issues promptly to ensure data accuracy and minimal downtime . Continuously explore opportunities for automation , performance improvements , and scalability enhancements in data workflows. Requirements: Strong programming skills in Python , PySpark , and SQL . Experience with Databricks and Spark -based data processing frameworks. Good understanding of ETL design principles , data modeling , and data architecture . Hands-on experience with workflow orchestration tools and version control systems (e.g., Git). Familiarity with cloud-based data platforms (AWS, Azure, or GCP) is an added advantage.
More at Colan Infotech