Padmi

Title Python Pyspark Developer

ChennaiPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Zorba AI

Opens the source posting on shine.com

Source description

About the role

View original

As a Python Data Engineer, you will be responsible for designing, developing, and maintaining robust and scalable data pipelines using Python and PySpark. Your key responsibilities will include: - Building and optimizing ETL/ELT processes for ingesting, transforming, and loading large volumes of structured and unstructured data. - Developing data processing solutions using Python libraries such as Pandas, NumPy, and PySpark. - Leveraging Databricks to implement and manage data engineering workflows and solve complex business problems. - Working with Azure Data Lake Storage (ADLS) and other Azure cloud services to manage and process data efficiently. - Writing complex SQL queries, stored procedures, and scripts for data extraction, transformation, validation, and reporting. - Collaborating with cross-functional teams including Data Analysts, Data Scientists, Architects, and Business Stakeholders to understand data requirements. - Monitoring, troubleshooting, and optimizing data pipelines for performance, scalability, and reliability. - Implementing data quality checks, governance standards, and security best practices. - Participating in code reviews and ensuring adherence to development standards and best practices. - Supporting CI/CD implementation and deployment activities following DevOps methodologies. To qualify for this role, you should have: - 58+ years of hands-on experience in Python development and data engineering. - Strong programming expertise in Python. - Experience with Python libraries such as Pandas, NumPy, and PySpark. - Strong knowledge of SQL and database concepts. - Hands-on experience with Databricks and Spark-based data processing. - Experience with Azure Cloud services and Azure Data Lake Storage (ADLS). - Solid understanding of ETL/ELT concepts and data warehousing principles. - Experience working with large-scale datasets and distributed computing frameworks. - Working knowledge of DevOps practices, CI/CD pipelines, Git, and deployment automation. - Strong analytical, problem-solving, and debugging skills. - Excellent communication and stakeholder management skills. Preferred skills for this role include experience with Azure Data Factory (ADF), Azure Synapse Analytics, or other Azure data services, knowledge of Delta Lake and Medallion Architecture, exposure to data governance, data security, and cloud-native architectures, and experience in Agile/Scrum development environments. You should hold a Bachelors or Masters degree in Computer Science, Information Technology, Engineering, or a related field. As a Python Data Engineer, you will be responsible for designing, developing, and maintaining robust and scalable data pipelines using Python and PySpark. Your key responsibilities will include: - Building and optimizing ETL/ELT processes for ingesting, transforming, and loading large volumes of structured and unstructured data. - Developing data processing solutions using Python libraries such as Pandas, NumPy, and PySpark. - Leveraging Databricks to implement and manage data engineering workflows and solve complex business problems. - Working with Azure Data Lake Storage (ADLS) and other Azure cloud services to manage and process data efficiently. - Writing complex SQL queries, stored procedures, and scripts for data extraction, transformation, validation, and reporting. - Collaborating with cross-functional teams including Data Analysts, Data Scientists, Architects, and Business Stakeholders to understand data requirements. - Monitoring, troubleshooting, and optimizing data pipelines for performance, scalability, and reliability. - Implementing data quality checks, governance standards, and security best practices. - Participating in code reviews and ensuring adherence to development standards and best practices. - Supporting CI/CD implementation and deployment activities following DevOps methodologies. To qualify for this role, you should have: - 58+ years of hands-on experience in Python development and data engineering. - Strong programming expertise in Python. - Experience with Python libraries such as Pandas, NumPy, and PySpark. - Strong knowledge of SQL and database concepts. - Hands-on experience with Databricks and Spark-based data processing. - Experience with Azure Cloud services and Azure Data Lake Storage (ADLS). - Solid understanding of ETL/ELT concepts and data warehousing principles. - Experience working with large-scale datasets and distributed computing frameworks. - Working knowledge of DevOps practices, CI/CD pipelines, Git, and deployment automation. - Strong analytical, problem-solving, and debugging skills. - Excellent communication and stakeholder management skills. Preferred skills for this role include experience with Azure Data Factory (ADF), Azure Synapse Analytics, or other Azure data services, knowledge of Delta Lake and Medallion Architecture, exposure to data governance, data security, and cloud-native architectures, and experience in Agi

One address, no account. We’ll tell you when matching roles go live.

More at Zorba AI

Related open roles

View all roles