Source description
About the role
Company: Celebal Technologies Location: Navi Mumbai Experience: 3-10 Years Role Overview We are looking for a skilled Data Engineer with strong experience in building scalable data pipelines using modern big data technologies. The ideal candidate should have hands-on experience with streaming + batch processing , along with deep knowledge of Databricks and Delta Lake. Key Responsibilities Design and build end-to-end ETL/ELT pipelines Ingest data from multiple sources like Kafka, databases, APIs, etc. Develop and optimize data pipelines using PySpark / Spark Work on both batch and real-time (streaming) data processing Implement Delta Lake architecture for reliable and scalable data storage Handle large-scale data (100GB1TB+) efficiently Perform data transformation, cleansing, and validation Optimize performance using partitioning, caching, and file formats Work with orchestration tools like Airflow for scheduling workflows Collaborate with Data Analysts and Business teams for requirements Required Skills (Must Have) Strong experience in PySpark / Apache Spark Hands-on with Databricks Good understanding of Delta Lake Experience with Apache Kafka Strong SQL skills (joins, CTEs, window functions, recursion basics) Experience in designing scalable data pipelines Understanding of data warehousing concepts Good to Have Experience with Microsoft Azure Knowledge of Apache Airflow Familiarity with Databricks Autoloader Experience with CI/CD tools like Jenkins Exposure to data modeling concepts What Were Looking For Someone who has worked on real production pipelines Strong problem-solving mindset Ability to handle large-scale data efficiently Clear understanding of streaming + batch architecture Interested candidate apply at chaity.mukherjee@celebaltech.com
More at Celebal Technologies