Source description
About the role
Python Pyspark JD
Job Summary: We are seeking a highly experienced Senior Python PySpark Developer with 10+ years of overall IT experience and strong expertise in building scalable data processing solutions. The ideal candidate will have deep hands-on experience in PySpark-based big data pipelines, performance optimization, and cloud-based data platforms, along with the ability to work in a coding-intensive environment.
Key Responsibilities: • Design, develop, and optimize large-scale data pipelines using Python and PySpark • Work with distributed data processing frameworks and handle high-volume datasets • Implement data transformation, cleansing, and aggregation logic using Spark • Collaborate with data engineers, architects, and business teams to deliver scalable solutions • Perform performance tuning and troubleshooting of Spark jobs • Write clean, efficient, and production-quality code following best practices • Participate in code reviews and provide technical guidance to junior team members • Support deployment, monitoring, and maintenance of data workflows
Required Skills: • 10+ years of overall IT experience • Strong hands-on expertise in Python and PySpark (mandatory) • Experience with Spark SQL, DataFrames, and RDD concepts • Strong understanding of distributed computing and big data architecture • Experience with ETL pipeline development and data processing frameworks • Knowledge of performance tuning and optimization of Spark jobs • Experience working with cloud platforms (AWS / Azure / GCP) • Strong SQL and data modeling skills • Familiarity with version control tools (Git) and CI/CD practices
More at Scadea