Source description
About the role
Role Overview: You will be responsible for designing, developing, and optimizing data pipelines using Python, PySpark, and SQL to support business intelligence, analytics, and data-driven decision-making. As a Python Data Engineer, you will work with Azure cloud services, Azure Data Lake Storage, and DevOps practices to create scalable solutions for managing and processing large datasets efficiently. Key Responsibilities: - Design, develop, and maintain robust and scalable data pipelines using Python and PySpark. - Build and optimize ETL/ELT processes for ingesting, transforming, and loading large volumes of structured and unstructured data. - Develop data processing solutions utilizing Python libraries such as Pandas, NumPy, and PySpark. - Utilize Databricks to implement and manage data engineering workflows and address complex business challenges. - Work with Azure Data Lake Storage (ADLS) and other Azure cloud services for efficient data management and processing. - Write complex SQL queries, stored procedures, and scripts for data extraction, transformation, validation, and reporting. - Collaborate with cross-functional teams to understand data requirements and contribute to data-driven decision-making. - Monitor, troubleshoot, and optimize data pipelines for performance, scalability, and reliability. - Implement data quality checks, governance standards, and security best practices. - Participate in code reviews, ensure adherence to development standards, and support CI/CD implementation following DevOps methodologies. Qualifications Required: - 58+ years of hands-on experience in Python development and data engineering. - Strong programming expertise in Python with experience in libraries such as Pandas, NumPy, and PySpark. - Proficient in SQL and database concepts. - Hands-on experience with Databricks and Spark-based data processing. - Familiarity with Azure Cloud services and Azure Data Lake Storage (ADLS). - Solid understanding of ETL/ELT concepts and data warehousing principles. - Experience working with large-scale datasets and distributed computing frameworks. - Knowledge of DevOps practices, CI/CD pipelines, Git, and deployment automation. - Strong analytical, problem-solving, and debugging skills. - Excellent communication and stakeholder management skills. Company Details: (Omit this section as there are no additional details of the company mentioned in the job description) Role Overview: You will be responsible for designing, developing, and optimizing data pipelines using Python, PySpark, and SQL to support business intelligence, analytics, and data-driven decision-making. As a Python Data Engineer, you will work with Azure cloud services, Azure Data Lake Storage, and DevOps practices to create scalable solutions for managing and processing large datasets efficiently. Key Responsibilities: - Design, develop, and maintain robust and scalable data pipelines using Python and PySpark. - Build and optimize ETL/ELT processes for ingesting, transforming, and loading large volumes of structured and unstructured data. - Develop data processing solutions utilizing Python libraries such as Pandas, NumPy, and PySpark. - Utilize Databricks to implement and manage data engineering workflows and address complex business challenges. - Work with Azure Data Lake Storage (ADLS) and other Azure cloud services for efficient data management and processing. - Write complex SQL queries, stored procedures, and scripts for data extraction, transformation, validation, and reporting. - Collaborate with cross-functional teams to understand data requirements and contribute to data-driven decision-making. - Monitor, troubleshoot, and optimize data pipelines for performance, scalability, and reliability. - Implement data quality checks, governance standards, and security best practices. - Participate in code reviews, ensure adherence to development standards, and support CI/CD implementation following DevOps methodologies. Qualifications Required: - 58+ years of hands-on experience in Python development and data engineering. - Strong programming expertise in Python with experience in libraries such as Pandas, NumPy, and PySpark. - Proficient in SQL and database concepts. - Hands-on experience with Databricks and Spark-based data processing. - Familiarity with Azure Cloud services and Azure Data Lake Storage (ADLS). - Solid understanding of ETL/ELT concepts and data warehousing principles. - Experience working with large-scale datasets and distributed computing frameworks. - Knowledge of DevOps practices, CI/CD pipelines, Git, and deployment automation. - Strong analytical, problem-solving, and debugging skills. - Excellent communication and stakeholder management skills. Company Details: (Omit this section as there are no additional details of the company mentioned in the job description)
More at Zorba AI