Source description
About the role
Data QA Engineer Location: Kolkata, India (Onsite/Hybrid) Experience: 24 years Employment Type: Full-time ABOUT THE ROLE We are hiring a Data QA Engineer for our data engineering team. This role is responsible for validating ETL pipelines and data platforms built on Azure Data Factory, Databricks, and Microsoft Fabric, along with the downstream tables, reports, and models they feed. The core responsibility is verifying data correctness, which is distinct from confirming that a pipeline executed without errors the two are frequently conflated, and this role exists to keep them separate. KEY RESPONSIBILITIES Design and execute test plans covering source-to-target validation and transformation logicfor ETL pipelinesWrite SQL and PySpark scripts to verify accuracy, completeness, and consistency in DeltaLake tablesPerform regression testing on every pipeline change; a successful pipeline run does notguarantee correct output, and validation must be independent of execution statusBuild and maintain reusable data quality checks (e.g., Great Expectations, dbt tests, orcustom PySpark frameworks) in place of one-off manual queriesReconcile data between source systems and target Lakehouse/Warehouse layers, andinvestigate root cause when discrepancies are foundValidate schema conformance, null/duplicate handling, referential integrity, and businessrule adherence across Bronze/Silver/Gold layersValidate Microsoft Fabric artifacts Lakehouses, Warehouses, and semantic models including DirectLake mode behavior and cross-domain data consistencyApply consistent QA methodology across tools; the underlying platform (Databricks, Fabric,or otherwise) should not change how rigorously data is validatedDocument test cases and defects with enough detail for engineers to act on them withoutrequiring additional clarificationWork directly with the data engineering team on requirement clarification and defectresolutionREQUIRED SKILLS Strong SQL, with the ability to write validation queries independentlyWorking proficiency in Python/PySpark for scripting data checksSolid understanding of ETL concepts: staging, transformations, incremental loads, SCDhandlingHands-on experience with Azure Data Factory and DatabricksWorking knowledge of Microsoft Fabric (Lakehouse, Warehouse, semantic models)Familiarity with Delta Lake and medallion architectureAbility to read transformation logic and determine expected outputGeneral data QA methodology that transfers across tools and platforms, not skills tied to asingle vendor stackExperience with defect tracking and structured test documentation (Jira, TestRail, orequivalent)PREFERRED QUALIFICATIONSExperience with a data quality framework (Great Expectations, Deequ, dbt tests)Familiarity with Unity Catalog and general data governance/lineage toolingExperience validating pipelines in a regulated domain (pharma, healthcare, finance) wherelineage and auditability are requirementsExposure to CI/CD for test automationAzure, Databricks, or Fabric certificationQUALIFICATIONS Bachelors degree in Computer Science, IT, or a related field 24 years of experience in data QA, data validation, or ETL testing; manual/UI testing experience without pipeline exposure does not meet this requirement CANDIDATE FIT This role requires the ability to determine why a data discrepancy occurred, not simply flag that one exists. Candidates whose QA background is primarily manual/UI testing and who are seeking to transition into data-focused work should not apply for this position; the required data depth is expected from day one. .