Source description
About the role
Role - Data Engineer Experience - 3-5 yrs Location - Bangalore About the Rol eThe Product Support Engineer will handle first-level technical investigation and mitigation for production issues by executing predefined command playbooks. You will work closely with engineering/SRE teams to keep customer-facing systems stable and healthy . We are looking for a Data Engineer II (SDE-2 ) to join our data team. The ideal candidate will be a play a key role to develop of high performant and scalable Data Lake-hous e, moving us toward a world of sub-minute data latency and unified batch/streaming compute. This is an engineering-heavy role where you will manage complex CDC flows, optimize distributed query engines and leverage AI to accelerate our development lifecycle .Technical Prioritie sReal-time CDC : Ownership of high-throughput ingestion from RDBMS to Lakehouse using Debeziu m, PeerD B .Lakehouse Architecture : Designing and optimizing table formats (Iceberg, Delta, Hud i) for both performance and storage efficiency .Unified Compute : Developing robust ETL/ELT frameworks in PySpar k and Flin k (handling both batch and streaming workloads) .Infrastructure & Ops : Managing data workloads on AWS (EMR, EKS, MSK, S3 ) and automating everything via Gitlab/Github Action s .Query & BI : Tuning Trin o or Clickhous e to power real-time dashboards in Metabas e, Superse t, and PowerB I .Requirement sExperience : 3–5 years in Data Engineering, specifically with distributed systems and cloud-native architectures .Coding : Expert-level Python/PySpar k and SQ L .Familiarity with Go/Java/Scal a is a plu sInfrastructure : Hands-on experience with AW S (S3, EKS, MSK) and Infrastructure-as-Code .Orchestration : Experience with Airflo w or Tempora l for complex workflow management .AI-Native : Proficiency in using AI tools (Claude, Codex, Copilo t) to write, test, and document code efficiently .Systems Thinking : Ability to explain the trade-offs between different storage formats and processing frameworks .Tech LeaderShip : Drive key tech initiatives by preparing TRD and actively involve in design reviews .Domain Modelling - Should be hands on in designing Domain models for OLAP like Fact, Dimension and types of SCD's and OBT pattern tables .Self Starter - Lead the team technically and bring in new ideas to contribute to the growth of the charter .Customer First - Interact with the Product & Key Stakeholders & help them by adding value to the business workflow with data & analytics .Our Tech Stac kIngestion : Debezium, PeerDB, Olak eStorage : Delta, Iceberg, Hudi (S3-based Lakehouse )Compute : PySpark, Flink, EMR, EK SStreaming : MSK (Kafka )Query Engines : Trino, Clickhous eOrchestration : Airflow, Tempora lDevOps : Gitlab, Github Actions, Terrafor mVisualization : Metabase, Superset, Tableau, PowerB I
More at Recro