Source description
About the role
Job Summary:We are seeking a highly experienced Senior Python PySpark Developer with 10+ years of overall IT experience and strong expertise in building scalable data processing solutions. The ideal candidate will have deep hands-on experience in PySpark-based big data pipelines, performance optimization, and cloud-based data platforms, along with the ability to work in a coding-intensive environment.
Key Responsibilities:- Design, develop, and optimize large-scale data pipelines using Python and PySpark
- Work with distributed data processing frameworks and handle high-volume datasets
- Implement data transformation, cleansing, and aggregation logic using Spark
- Collaborate with data engineers, architects, and business teams to deliver scalable solutions
- Perform performance tuning and troubleshooting of Spark jobs
- Write clean, efficient, and production-quality code following best practices
- Participate in code reviews and provide technical guidance to junior team members
- Support deployment, monitoring, and maintenance of data workflows Required Skills:- 10+ years of overall IT experience
- Strong hands-on expertise in Python and PySpark (mandatory)
- Experience with Spark SQL, DataFrames, and RDD concepts
- Strong understanding of distributed computing and big data architecture
- Experience with ETL pipeline development and data processing frameworks
- Knowledge of performance tuning and optimization of Spark jobs
- Experience working with cloud platforms (AWS / Azure / GCP)
- Strong SQL and data modeling skills
- Familiarity with version control tools (Git) and CI/CD practices
More at Scadea