Padmi

Python Pyspark Developer

Delhi NCRPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at Zorba AI

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: You will be responsible for designing, developing, and optimizing data pipelines using Python, PySpark, and SQL to support business intelligence, analytics, and data-driven decision-making. As a Python Data Engineer, you will work with Azure cloud services, Azure Data Lake Storage, and DevOps practices to create scalable solutions for managing and processing large datasets efficiently. Key Responsibilities: - Design, develop, and maintain robust and scalable data pipelines using Python and PySpark. - Build and optimize ETL/ELT processes for ingesting, transforming, and loading large volumes of structured and unstructured data. - Develop data processing solutions utilizing Python libraries such as Pandas, NumPy, and PySpark. - Utilize Databricks to implement and manage data engineering workflows and address complex business challenges. - Work with Azure Data Lake Storage (ADLS) and other Azure cloud services for efficient data management and processing. - Write complex SQL queries, stored procedures, and scripts for data extraction, transformation, validation, and reporting. - Collaborate with cross-functional teams to understand data requirements and contribute to data-driven decision-making. - Monitor, troubleshoot, and optimize data pipelines for performance, scalability, and reliability. - Implement data quality checks, governance standards, and security best practices. - Participate in code reviews, ensure adherence to development standards, and support CI/CD implementation following DevOps methodologies. Qualifications Required: - 58+ years of hands-on experience in Python development and data engineering. - Strong programming expertise in Python with experience in libraries such as Pandas, NumPy, and PySpark. - Proficient in SQL and database concepts. - Hands-on experience with Databricks and Spark-based data processing. - Familiarity with Azure Cloud services and Azure Data Lake Storage (ADLS). - Solid understanding of ETL/ELT concepts and data warehousing principles. - Experience working with large-scale datasets and distributed computing frameworks. - Knowledge of DevOps practices, CI/CD pipelines, Git, and deployment automation. - Strong analytical, problem-solving, and debugging skills. - Excellent communication and stakeholder management skills. Company Details: (Omit this section as there are no additional details of the company mentioned in the job description) Role Overview: You will be responsible for designing, developing, and optimizing data pipelines using Python, PySpark, and SQL to support business intelligence, analytics, and data-driven decision-making. As a Python Data Engineer, you will work with Azure cloud services, Azure Data Lake Storage, and DevOps practices to create scalable solutions for managing and processing large datasets efficiently. Key Responsibilities: - Design, develop, and maintain robust and scalable data pipelines using Python and PySpark. - Build and optimize ETL/ELT processes for ingesting, transforming, and loading large volumes of structured and unstructured data. - Develop data processing solutions utilizing Python libraries such as Pandas, NumPy, and PySpark. - Utilize Databricks to implement and manage data engineering workflows and address complex business challenges. - Work with Azure Data Lake Storage (ADLS) and other Azure cloud services for efficient data management and processing. - Write complex SQL queries, stored procedures, and scripts for data extraction, transformation, validation, and reporting. - Collaborate with cross-functional teams to understand data requirements and contribute to data-driven decision-making. - Monitor, troubleshoot, and optimize data pipelines for performance, scalability, and reliability. - Implement data quality checks, governance standards, and security best practices. - Participate in code reviews, ensure adherence to development standards, and support CI/CD implementation following DevOps methodologies. Qualifications Required: - 58+ years of hands-on experience in Python development and data engineering. - Strong programming expertise in Python with experience in libraries such as Pandas, NumPy, and PySpark. - Proficient in SQL and database concepts. - Hands-on experience with Databricks and Spark-based data processing. - Familiarity with Azure Cloud services and Azure Data Lake Storage (ADLS). - Solid understanding of ETL/ELT concepts and data warehousing principles. - Experience working with large-scale datasets and distributed computing frameworks. - Knowledge of DevOps practices, CI/CD pipelines, Git, and deployment automation. - Strong analytical, problem-solving, and debugging skills. - Excellent communication and stakeholder management skills. Company Details: (Omit this section as there are no additional details of the company mentioned in the job description)

One address, no account. We’ll tell you when matching roles go live.

More at Zorba AI

Related open roles

View all roles