Source description
About the role
Responsiblities Build batch-seed and event-tail ingestion per source system, including seed tail watermark handoff, idempotent upserts, and dedup ledgers Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints, retry/backoff, DLQ), and full observability Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation financial is zero-tolerance Build identity matching combining deterministic rules with probabilistic scoring and confidence bands; deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data Author and maintain source canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution What Were Looking For 5+ years building production data pipelines at scale Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns Strong SQL and solid ETL fundamentals Java and/or Python in production Medallion/lakehouse layering, CDC, watermark/checkpoint patterns, and batch-stream hand-off Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation Entity resolution / MDM exposure: record matching, dedup, survivorship via commercial tools (Informatica MDM, Reltio) or custom builds Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data, config-as-code (YAML/JSON, Git) Bonus Points Probabilistic record linkage at depth blocking/candidate generation, scoring models, threshold calibration (expected at senior level) Schema registry experience (Avro/Protobuf) Extracting from mainframe or older RDBMS sources with limited CDC support Financial reconciliation in finance-adjacent domains Benefits administration or healthcare domain knowledge Responsiblities Build batch-seed and event-tail ingestion per source system, including seed tail watermark handoff, idempotent upserts, and dedup ledgers Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints, retry/backoff, DLQ), and full observability Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine covering count, record, and financial reconciliation financial is zero-tolerance Build identity matching combining deterministic rules with probabilistic scoring and confidence bands; deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds with labelled data Author and maintain source canonical structural mappings and value crosswalks (e.g., collapsing 1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible schema evolution What Were Looking For 5+ years building production data pipelines at scale Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns Strong SQL and solid ETL fundamentals Java and/or Python in production Medallion/lakehouse layering, CDC, watermark/checkpoint patterns, and batch-stream hand-off Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation Entity resolution / MDM exposure: record matching, dedup, survivorship via commercial tools (Informatica MDM, Reltio) or custom builds Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data, config-as-code (YAML/JSON, Git) Bonus Points Probabilistic record linkage at depth blocking/candidate generation, scoring models, threshold calibration (expected at senior level) Schema registry experience (Avro/Protobuf) Extracting from mainframe or older RDBMS sources with limited CDC support Financial reconciliation in finance-adjacent domains Benefits administration or healthcare domain knowledge
More at SE MENTOR SOLUTIONS