Padmi

Data Engineer - ETL/Python

MumbaiPosted 2 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at TalentSyker

Opens the source posting on shine.com

Source description

About the role

View original

Responsibilities : - Design, build, and manage scalable data pipelines, ensuring user event data is reliably ingested into the data warehouse. - Develop and maintain canonical datasets to track key product metrics such as user growth, engagement, retention, and revenue. - Collaborate with Infrastructure, Data Science, Product, Marketing, Finance, and Research teams to understand data needs and deliver effective solutions. - Implement robust, fault-tolerant systems for data ingestion, transformation, and processing. - Participate actively in data architecture and engineering decisions, contributing best practices and long-term scalability thinking. - Ensure data security, integrity, and compliance in line with company policies and industry standards. - Monitor pipeline health, troubleshoot failures, and continuously improve reliability and performance. Requirements : - 3-5 years of professional experience working as a data engineer or in a similar role. - Proficiency in at least one data engineering programming language, such as Python, Scala, or Java. - Experience with distributed data processing frameworks and technologies such as Hadoop, Flink, and distributed storage systems (e. g., HDFS). - Strong expertise with ETL orchestration tools, such as Apache Airflow. - Solid understanding of Apache Spark, with the ability to write, debug, and optimize Spark jobs. - Experience designing and maintaining data pipelines for analytics, reporting, or ML use cases. - Strong problem-solving skills and the ability to work across teams with varied data requirements. Desired Skills : - Hands-on experience working with Databricks in production environments. - Familiarity with the GCP data stack, including Pub/Sub, Dataflow, BigQuery, and Google Cloud Storage (GCS). - Exposure to data quality frameworks, data validation, or schema management tools. - Understanding of analytics use cases, experimentation, or ML data workflows. Responsibilities : - Design, build, and manage scalable data pipelines, ensuring user event data is reliably ingested into the data warehouse. - Develop and maintain canonical datasets to track key product metrics such as user growth, engagement, retention, and revenue. - Collaborate with Infrastructure, Data Science, Product, Marketing, Finance, and Research teams to understand data needs and deliver effective solutions. - Implement robust, fault-tolerant systems for data ingestion, transformation, and processing. - Participate actively in data architecture and engineering decisions, contributing best practices and long-term scalability thinking. - Ensure data security, integrity, and compliance in line with company policies and industry standards. - Monitor pipeline health, troubleshoot failures, and continuously improve reliability and performance. Requirements : - 3-5 years of professional experience working as a data engineer or in a similar role. - Proficiency in at least one data engineering programming language, such as Python, Scala, or Java. - Experience with distributed data processing frameworks and technologies such as Hadoop, Flink, and distributed storage systems (e. g., HDFS). - Strong expertise with ETL orchestration tools, such as Apache Airflow. - Solid understanding of Apache Spark, with the ability to write, debug, and optimize Spark jobs. - Experience designing and maintaining data pipelines for analytics, reporting, or ML use cases. - Strong problem-solving skills and the ability to work across teams with varied data requirements. Desired Skills : - Hands-on experience working with Databricks in production environments. - Familiarity with the GCP data stack, including Pub/Sub, Dataflow, BigQuery, and Google Cloud Storage (GCS). - Exposure to data quality frameworks, data validation, or schema management tools. - Understanding of analytics use cases, experimentation, or ML data workflows.

One address, no account. We’ll tell you when matching roles go live.

More at TalentSyker

Related open roles

View all roles