Padmi

Python, Pyspark, SQL

IndiaPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Cognizant

Opens the source posting on shine.com

Source description

About the role

View original

Skill: Python, Pyspark, SQL Exp: 6 to 12 years Location: Pune We are seeking a highly skilled Python / PySpark / SQL Developer to design, develop, and optimize large-scale data processing pipelines. The ideal candidate will have strong expertise in Python programming, PySpark for distributed data processing, and SQL for data querying and transformation. You will work closely with data engineers, analysts, and business stakeholders to deliver efficient, scalable, and reliable data solutions. Key Responsibilities Design, develop, and maintain ETL/ELT pipelines using PySpark and Python.Write optimized SQL queries for data extraction, transformation, and loading.Work with big data platforms (e.g., Hadoop, Databricks, AWS EMR, Azure Synapse, or GCP Dataproc).Implement data quality checks, validation, and error handling in pipelines.Optimize PySpark jobs for performance and scalability.Collaborate with cross-functional teams to understand data requirements and deliver solutions.Maintain documentation for data flows, transformations, and processes.Ensure data security, compliance, and governance standards are met.Troubleshoot and debug data processing issues in production environments. Required Skills & Qualifications 3+ years of experience in Python development.Strong hands-on experience with PySpark (RDDs, DataFrames, Spark SQL).Proficiency in SQL (complex joins, window functions, CTEs, performance tuning).Experience with big data ecosystems (HDFS, Hive, Delta Lake, etc.).Familiarity with cloud data platforms (AWS, Azure, or GCP).Strong understanding of data modeling and ETL best practices.Experience with version control (Git) and CI/CD pipelines.Knowledge of performance tuning for Spark and SQL queries.Excellent problem-solving and communication skills. Preferred Skills Experience with Airflow, Luigi, or other workflow orchestration tools.Knowledge of NoSQL databases (Cassandra, MongoDB, etc.).Familiarity with containerization (Docker, Kubernetes).Exposure to machine learning pipelines in Spark.Understanding of data warehousing concepts (Snowflake, Redshift, BigQuery). Skill: Python, Pyspark, SQL Exp: 6 to 12 years Location: Pune We are seeking a highly skilled Python / PySpark / SQL Developer to design, develop, and optimize large-scale data processing pipelines. The ideal candidate will have strong expertise in Python programming, PySpark for distributed data processing, and SQL for data querying and transformation. You will work closely with data engineers, analysts, and business stakeholders to deliver efficient, scalable, and reliable data solutions. Key Responsibilities Design, develop, and maintain ETL/ELT pipelines using PySpark and Python.Write optimized SQL queries for data extraction, transformation, and loading.Work with big data platforms (e.g., Hadoop, Databricks, AWS EMR, Azure Synapse, or GCP Dataproc).Implement data quality checks, validation, and error handling in pipelines.Optimize PySpark jobs for performance and scalability.Collaborate with cross-functional teams to understand data requirements and deliver solutions.Maintain documentation for data flows, transformations, and processes.Ensure data security, compliance, and governance standards are met.Troubleshoot and debug data processing issues in production environments. Required Skills & Qualifications 3+ years of experience in Python development.Strong hands-on experience with PySpark (RDDs, DataFrames, Spark SQL).Proficiency in SQL (complex joins, window functions, CTEs, performance tuning).Experience with big data ecosystems (HDFS, Hive, Delta Lake, etc.).Familiarity with cloud data platforms (AWS, Azure, or GCP).Strong understanding of data modeling and ETL best practices.Experience with version control (Git) and CI/CD pipelines.Knowledge of performance tuning for Spark and SQL queries.Excellent problem-solving and communication skills. Preferred Skills Experience with Airflow, Luigi, or other workflow orchestration tools.Knowledge of NoSQL databases (Cassandra, MongoDB, etc.).Familiarity with containerization (Docker, Kubernetes).Exposure to machine learning pipelines in Spark.Understanding of data warehousing concepts (Snowflake, Redshift, BigQuery).

One address, no account. We’ll tell you when matching roles go live.

More at Cognizant

Related open roles

View all roles