Source description
About the role
Job Title: Data Engineer GCP, PySpark & ScalaExperience 6+ Years Job Summary We are looking for a skilled Data Engineer with strong expertise in GCP, PySpark, and Scala to design, develop, deploy, and optimize scalable data pipelines. The ideal candidate should have hands-on experience building batch ETL workflows, implementing Medallion Architecture, and working with GCP data services. The role requires strong problem-solving skills and experience in performance optimization, workflow orchestration, and data quality. Key Responsibilities Design, develop, deploy, monitor, and optimize batch ETL/data pipelines using Scala, PySpark, and GCP services. Build and maintain scalable data processing workflows using Google Cloud Platform technologies such as BigQuery, Dataproc, and Cloud Storage. Develop and manage Airflow DAGs, ensuring proper dependency handling, scheduling, retry mechanisms, and idempotent execution. Implement Medallion Architecture (Bronze, Silver, Gold) with robust data quality checks and data summarization processes. Optimize ETL jobs and Spark applications for performance, scalability, and cost efficiency. Monitor production pipelines, create observability dashboards, and develop runbooks for operational support and continuous optimization. Troubleshoot production issues and collaborate with cross-functional teams to deliver reliable data solutions. Follow coding standards, best practices, and CI/CD processes for data engineering solutions. Support data governance initiatives and ensure compliance with enterprise data management standards. Required Skills 6+ years of experience in Data Engineering. Strong hands-on experience with Scala and PySpark. Expertise in Google Cloud Platform (GCP) services, including BigQuery, Dataproc, Cloud Storage, and related data services. Experience developing, deploying, and optimizing batch ETL workflows. Strong knowledge of Apache Airflow, including workflow orchestration, dependency management, retries, and idempotent execution. Experience implementing Medallion Architecture and data quality frameworks. Strong SQL skills and experience with large-scale data processing. Experience in performance tuning, monitoring, observability, and operational support. Good understanding of distributed data processing and Spark optimization techniques. Preferred Skills Exposure to data governance and metadata management. Experience with CI/CD pipelines and Git. Familiarity with Agile/Scrum development methodologies. Knowledge of data security and cloud best practices. Skills: gcp,pyspark,airproc,etl,scala Job Title: Data Engineer GCP, PySpark & ScalaExperience 6+ Years Job Summary We are looking for a skilled Data Engineer with strong expertise in GCP, PySpark, and Scala to design, develop, deploy, and optimize scalable data pipelines. The ideal candidate should have hands-on experience building batch ETL workflows, implementing Medallion Architecture, and working with GCP data services. The role requires strong problem-solving skills and experience in performance optimization, workflow orchestration, and data quality. Key Responsibilities Design, develop, deploy, monitor, and optimize batch ETL/data pipelines using Scala, PySpark, and GCP services. Build and maintain scalable data processing workflows using Google Cloud Platform technologies such as BigQuery, Dataproc, and Cloud Storage. Develop and manage Airflow DAGs, ensuring proper dependency handling, scheduling, retry mechanisms, and idempotent execution. Implement Medallion Architecture (Bronze, Silver, Gold) with robust data quality checks and data summarization processes. Optimize ETL jobs and Spark applications for performance, scalability, and cost efficiency. Monitor production pipelines, create observability dashboards, and develop runbooks for operational support and continuous optimization. Troubleshoot production issues and collaborate with cross-functional teams to deliver reliable data solutions. Follow coding standards, best practices, and CI/CD processes for data engineering solutions. Support data governance initiatives and ensure compliance with enterprise data management standards. Required Skills 6+ years of experience in Data Engineering. Strong hands-on experience with Scala and PySpark. Expertise in Google Cloud Platform (GCP) services, including BigQuery, Dataproc, Cloud Storage, and related data services. Experience developing, deploying, and optimizing batch ETL workflows. Strong knowledge of Apache Airflow, including workflow orchestration, dependency management, retries, and idempotent execution. Experience implementing Medallion Architecture and data quality frameworks. Strong SQL skills and experience with large-scale data processing. Experience in performance tuning, monitoring, observability, and operational support. Good understanding of distributed data processing and Spark optimization techniques. Preferred Skills Exposure to data governance and metadata ma
More at Zorba AI