Source description
About the role
Responsibilities : - Design, build, and manage scalable data pipelines, ensuring user event data is reliably ingested into the data warehouse. - Develop and maintain canonical datasets to track key product metrics such as user growth, engagement, retention, and revenue. - Collaborate with Infrastructure, Data Science, Product, Marketing, Finance, and Research teams to understand data needs and deliver effective solutions. - Implement robust, fault-tolerant systems for data ingestion, transformation, and processing. - Participate actively in data architecture and engineering decisions, contributing best practices and long-term scalability thinking. - Ensure data security, integrity, and compliance in line with company policies and industry standards. - Monitor pipeline health, troubleshoot failures, and continuously improve reliability and performance. Requirements : - 3-5 years of professional experience working as a data engineer or in a similar role. - Proficiency in at least one data engineering programming language, such as Python, Scala, or Java. - Experience with distributed data processing frameworks and technologies such as Hadoop, Flink, and distributed storage systems (e. g., HDFS). - Strong expertise with ETL orchestration tools, such as Apache Airflow. - Solid understanding of Apache Spark, with the ability to write, debug, and optimize Spark jobs. - Experience designing and maintaining data pipelines for analytics, reporting, or ML use cases. - Strong problem-solving skills and the ability to work across teams with varied data requirements. Desired Skills : - Hands-on experience working with Databricks in production environments. - Familiarity with the GCP data stack, including Pub/Sub, Dataflow, BigQuery, and Google Cloud Storage (GCS). - Exposure to data quality frameworks, data validation, or schema management tools. - Understanding of analytics use cases, experimentation, or ML data workflows. Responsibilities : - Design, build, and manage scalable data pipelines, ensuring user event data is reliably ingested into the data warehouse. - Develop and maintain canonical datasets to track key product metrics such as user growth, engagement, retention, and revenue. - Collaborate with Infrastructure, Data Science, Product, Marketing, Finance, and Research teams to understand data needs and deliver effective solutions. - Implement robust, fault-tolerant systems for data ingestion, transformation, and processing. - Participate actively in data architecture and engineering decisions, contributing best practices and long-term scalability thinking. - Ensure data security, integrity, and compliance in line with company policies and industry standards. - Monitor pipeline health, troubleshoot failures, and continuously improve reliability and performance. Requirements : - 3-5 years of professional experience working as a data engineer or in a similar role. - Proficiency in at least one data engineering programming language, such as Python, Scala, or Java. - Experience with distributed data processing frameworks and technologies such as Hadoop, Flink, and distributed storage systems (e. g., HDFS). - Strong expertise with ETL orchestration tools, such as Apache Airflow. - Solid understanding of Apache Spark, with the ability to write, debug, and optimize Spark jobs. - Experience designing and maintaining data pipelines for analytics, reporting, or ML use cases. - Strong problem-solving skills and the ability to work across teams with varied data requirements. Desired Skills : - Hands-on experience working with Databricks in production environments. - Familiarity with the GCP data stack, including Pub/Sub, Dataflow, BigQuery, and Google Cloud Storage (GCS). - Exposure to data quality frameworks, data validation, or schema management tools. - Understanding of analytics use cases, experimentation, or ML data workflows.
More at TalentSyker