Padmi
Citi logo
Citi

credit cards · retail banking

Python and PySpark Developer

Delhi NCRPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Citi

Opens the source posting on shine.com

Source description

About the role

View original

As a Python / PySpark Developer at our company, your role will involve supporting the development and maintenance of scalable data processing solutions. You will have the opportunity to work with senior engineers and collaborate with data teams to build reliable data pipelines and contribute to analytics and reporting solutions. Key Responsibilities: - Assist in developing and maintaining data pipelines using Python and PySpark - Support ETL/ELT workflows for batch data processing - Write clean, readable, and well-structured Python code following best practices - Perform basic data transformations, aggregations, and validations - Debug and troubleshoot pipeline issues with guidance from senior developers - Work with structured and semi-structured data formats (CSV, JSON, Parquet, etc.) - Assist in integrating data from databases, APIs, and cloud storage systems - Help ensure data quality and consistency within pipelines - Support migration of legacy scripts to modern data platforms - Collaborate with team members on development tasks and code reviews - Participate in knowledge-sharing and training sessions - Learn and adopt new tools, frameworks, and best practices - Assist in documenting data workflows and technical processes Required Skills & Qualifications: Technical Skills: - Basic to intermediate proficiency in Python - 4-7 years of experience - Exposure to Apache Spark / PySpark (internship or project experience is acceptable) - Understanding of fundamental programming and data structures - Basic knowledge of SQL and relational databases - Familiarity with data processing concepts and ETL fundamentals - Awareness of Linux/Unix command line is a plus Engineering Fundamentals: - Understanding of coding best practices and version control (Git) - Basic debugging and problem-solving skills - Exposure to unit testing concepts is a plus Nice to Have (Preferred Skills): - Exposure to big data tools (Hive, Hadoop ecosystem, or similar) - Familiarity with cloud platforms (AWS / Azure / GCP) - Basic knowledge of job orchestration tools (Airflow, etc.) - Understanding of data pipelines and workflow lifecycle - Academic or project experience with data engineering or analytics In addition to the technical skills and qualifications required, ideal candidate traits include a strong willingness to learn and grow in a fast-paced environment, good analytical and problem-solving skills, effective communication and teamwork abilities, as well as attention to detail and commitment to quality. As a Python / PySpark Developer at our company, your role will involve supporting the development and maintenance of scalable data processing solutions. You will have the opportunity to work with senior engineers and collaborate with data teams to build reliable data pipelines and contribute to analytics and reporting solutions. Key Responsibilities: - Assist in developing and maintaining data pipelines using Python and PySpark - Support ETL/ELT workflows for batch data processing - Write clean, readable, and well-structured Python code following best practices - Perform basic data transformations, aggregations, and validations - Debug and troubleshoot pipeline issues with guidance from senior developers - Work with structured and semi-structured data formats (CSV, JSON, Parquet, etc.) - Assist in integrating data from databases, APIs, and cloud storage systems - Help ensure data quality and consistency within pipelines - Support migration of legacy scripts to modern data platforms - Collaborate with team members on development tasks and code reviews - Participate in knowledge-sharing and training sessions - Learn and adopt new tools, frameworks, and best practices - Assist in documenting data workflows and technical processes Required Skills & Qualifications: Technical Skills: - Basic to intermediate proficiency in Python - 4-7 years of experience - Exposure to Apache Spark / PySpark (internship or project experience is acceptable) - Understanding of fundamental programming and data structures - Basic knowledge of SQL and relational databases - Familiarity with data processing concepts and ETL fundamentals - Awareness of Linux/Unix command line is a plus Engineering Fundamentals: - Understanding of coding best practices and version control (Git) - Basic debugging and problem-solving skills - Exposure to unit testing concepts is a plus Nice to Have (Preferred Skills): - Exposure to big data tools (Hive, Hadoop ecosystem, or similar) - Familiarity with cloud platforms (AWS / Azure / GCP) - Basic knowledge of job orchestration tools (Airflow, etc.) - Understanding of data pipelines and workflow lifecycle - Academic or project experience with data engineering or analytics In addition to the technical skills and qualifications required, ideal candidate traits include a strong willingness to learn and grow in a fast-paced environment, good analytical and problem-solving skills, ef

One address, no account. We’ll tell you when matching roles go live.

More at Citi

Related open roles

View all roles