Padmi

Py Spark Data Engineer

BangalorePosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at Ensoft Consulting Pte Ltd

Opens the source posting on shine.com

Source description

About the role

View original

Pyspark Data Engineer:Hands-on expertise in designing, building, and maintaining Apache Spark pipelines in production environments. Proven experience building and scaling data ingestion frameworks that integrate data from multiple source systems, with a focus on reliability, reusability, and scalability. Deep understanding of Spark architecture (driver/executors, DAG, partitioning, shuffles, caching, cluster resource management) and experience operating pipelines at scale, including data transformations on datasets 500 GB+. Strong understanding of Oracle SQL and HDFS, including handling file formats and applying appropriate data cleansing, normalization, and formatting to produce curated output datasets. Ability to write Python, Pyspark, and shell scripts to process, transform, and automate data workflows. The Candidate should be good in writing application programs and automation manual data processing steps using python. PySpark Developer / Senior Data EngineerSkills Strong hands-on experience in PySpark, Python, and SQL.Experience designing and optimizing Spark-based ETL/ELT pipelines and data processing jobs.Strong understanding toa BigQuery.Strong understanding of data quality, governance, observability, and performance tuning.Good collaboration, debugging, and Agile delivery skills. Experience Bachelor's or Master's degree plus 6+ years of data engineering experience with strong PySpark expertise. Pyspark Data Engineer:Hands-on expertise in designing, building, and maintaining Apache Spark pipelines in production environments. Proven experience building and scaling data ingestion frameworks that integrate data from multiple source systems, with a focus on reliability, reusability, and scalability. Deep understanding of Spark architecture (driver/executors, DAG, partitioning, shuffles, caching, cluster resource management) and experience operating pipelines at scale, including data transformations on datasets 500 GB+. Strong understanding of Oracle SQL and HDFS, including handling file formats and applying appropriate data cleansing, normalization, and formatting to produce curated output datasets. Ability to write Python, Pyspark, and shell scripts to process, transform, and automate data workflows. The Candidate should be good in writing application programs and automation manual data processing steps using python. PySpark Developer / Senior Data EngineerSkills Strong hands-on experience in PySpark, Python, and SQL.Experience designing and optimizing Spark-based ETL/ELT pipelines and data processing jobs.Strong understanding toa BigQuery.Strong understanding of data quality, governance, observability, and performance tuning.Good collaboration, debugging, and Agile delivery skills. Experience Bachelor's or Master's degree plus 6+ years of data engineering experience with strong PySpark expertise.

One address, no account. We’ll tell you when matching roles go live.

More at Ensoft Consulting Pte Ltd

Related open roles

View all roles
Py Spark Data Engineer at Ensoft Consulting Pte Ltd · Padmi