Source description
About the role
Role: Data Engineer II (SDE-2) Experience: 3–5 Years Joining: IMMEDIATE About the Role We are looking for a Data Engineer II (SDE-2) to join our Data Platform team and build scalable, high-performance data infrastructure powering real-time analytics and business-critical applications. In this role, you will design and develop modern Lakehouse architectures , build low-latency data pipelines , and optimize distributed data processing systems on AWS . You will work closely with Product, Engineering, and Analytics teams to deliver reliable, scalable, and efficient data solutions. Key Responsibilities Design, build, and maintain scalable batch and real-time data pipelines. Develop and manage Change Data Capture (CDC) pipelines using technologies like Debezium, PeerDB , or similar. Design and optimize Lakehouse architectures using Apache Iceberg, Delta Lake, or Apache Hudi . Build robust ETL/ELT pipelines using PySpark and Apache Flink . Manage and optimize data infrastructure on AWS (S3, EMR, EKS, MSK) . Tune distributed query engines such as Trino or ClickHouse for high-performance analytics. Design dimensional data models including Fact Tables, Dimension Tables, SCDs, and OBTs . Automate deployments and infrastructure using GitLab/GitHub Actions and Terraform . Drive technical design discussions, architecture reviews, and documentation. Collaborate with Product and Business teams to deliver data solutions that enable analytics and decision-making. Required Skills 3–5 years of experience in Data Engineering building scalable data platforms. Strong programming skills in Python (PySpark) and SQL . Hands-on experience with Apache Spark and distributed data processing. Experience building batch and streaming data pipelines. Strong knowledge of AWS services such as S3, EMR, EKS, and MSK . Experience with workflow orchestration tools like Airflow or Temporal . Understanding of modern Lakehouse technologies ( Iceberg, Delta Lake, or Hudi ). Experience with distributed query engines such as Trino or ClickHouse . Good understanding of data modeling concepts including Fact/Dimension modeling and Slowly Changing Dimensions (SCDs) . Strong problem-solving skills and ability to work in a fast-paced product engineering environment. Good to Have Experience with Kafka , Debezium , or PeerDB . Exposure to Terraform , GitLab CI/CD , or GitHub Actions . Programming experience in Go, Java, or Scala . Experience using AI-assisted development tools such as GitHub Copilot, Claude, or Codex . Tech Stack Languages: Python, SQL (Go/Java/Scala – good to have) Data Processing: PySpark, Apache Flink Streaming & CDC: Kafka (MSK), Debezium, PeerDB Storage: Apache Iceberg, Delta Lake, Apache Hudi Cloud: AWS (S3, EMR, EKS, MSK) Query Engines: Trino, ClickHouse Orchestration: Airflow, Temporal DevOps: GitLab CI/CD, GitHub Actions, Terraform BI & Visualization: Metabase, Superset, Tableau, Power BI
More at Recro