Padmi

Python Pyspark Developer

IndiaPosted 1 month ago
Software engineeringJuniorFull Time; Regular
Apply at Viraaj HR Solutions Private Limited

Opens the source posting on shine.com

Source description

About the role

View original

About The Opportunity A fast-scaling technology services firm operating in the Data Engineering and Cloud Analytics space, we partner with enterprises to build scalable data pipelines, automate ETL workflows, and unlock real-time insights from structured and unstructured datasets. Our Python PySpark developers architect and deploy high-performance data solutions on cloud platformsenabling faster decision-making, predictive modeling, and regulatory compliance for global clients. Role & Responsibilities Design, develop, and optimize PySpark ETL pipelines for large-scale batch and streaming data processing.Collaborate with data scientists and analysts to ingest, transform, and model data for ML training and reporting use cases.Implement data quality checks, error handling, and logging frameworks within Spark jobs for production-grade reliability.Integrate PySpark workflows with cloud data lakes (AWS S3, Azure Data Lake), metastores (Hive, Glue), and orchestration tools (Airflow, Luigi).Performance-tune Spark applicationspartitioning, caching, broadcasting, and resource allocationto reduce job runtimes and costs.Write clean, modular, testable Python code with unit/integration tests and document architecture decisions for team scalability. Skills & Qualifications Must-Have PySparkPythonApache SparkETL DevelopmentSQLData Lake ArchitectureAWS S3Apache Airflow Preferred Azure Data LakeDatabricksDelta Lake Benefits & Culture Highlights On-site collaborative workspace in major Indian tech hubs with modern infrastructure and R&D labs.Opportunities to upskill in cloud-native data platforms and GenAI-powered data pipelines.Fast-track career growth with cross-functional exposure to data science, ML engineering, and cloud architecture teams. Skills: data,python,lake,spark,cloud About The Opportunity A fast-scaling technology services firm operating in the Data Engineering and Cloud Analytics space, we partner with enterprises to build scalable data pipelines, automate ETL workflows, and unlock real-time insights from structured and unstructured datasets. Our Python PySpark developers architect and deploy high-performance data solutions on cloud platformsenabling faster decision-making, predictive modeling, and regulatory compliance for global clients. Role & Responsibilities Design, develop, and optimize PySpark ETL pipelines for large-scale batch and streaming data processing.Collaborate with data scientists and analysts to ingest, transform, and model data for ML training and reporting use cases.Implement data quality checks, error handling, and logging frameworks within Spark jobs for production-grade reliability.Integrate PySpark workflows with cloud data lakes (AWS S3, Azure Data Lake), metastores (Hive, Glue), and orchestration tools (Airflow, Luigi).Performance-tune Spark applicationspartitioning, caching, broadcasting, and resource allocationto reduce job runtimes and costs.Write clean, modular, testable Python code with unit/integration tests and document architecture decisions for team scalability. Skills & Qualifications Must-Have PySparkPythonApache SparkETL DevelopmentSQLData Lake ArchitectureAWS S3Apache Airflow Preferred Azure Data LakeDatabricksDelta Lake Benefits & Culture Highlights On-site collaborative workspace in major Indian tech hubs with modern infrastructure and R&D labs.Opportunities to upskill in cloud-native data platforms and GenAI-powered data pipelines.Fast-track career growth with cross-functional exposure to data science, ML engineering, and cloud architecture teams. Skills: data,python,lake,spark,cloud

One address, no account. We’ll tell you when matching roles go live.

More at Viraaj HR Solutions Private Limited

Related open roles

View all roles