Source description
About the role
Experience: 58+ Years Job Summary We are looking for a strong Data Validation / ETL Testing Engineer with hands-on experience in validating complex data pipelines across cloud and big data platforms. The ideal candidate should have expertise in Python, AWS services, Databricks, PySpark, Redshift , and ETL validation frameworks to ensure high-quality, accurate, and reliable data solutions. Required Skills Strong experience in Data Validation and ETL Testing . Hands-on expertise in Python for automation and data validation. Experience with AWS Services , including: S3 Glue Redshift Lambda Strong knowledge of Databricks and PySpark . Experience validating large-scale ETL/ELT data pipelines. Expertise in writing complex SQL queries for data validation and reconciliation. Experience with ETL validation frameworks and automation testing. Understanding of data warehousing concepts and dimensional modeling. Knowledge of data quality, data profiling, and reconciliation techniques. Familiarity with Agile/Scrum methodologies. Strong analytical, debugging, and problem-solving skills. Roles & Responsibilities Validate end-to-end ETL/ELT data pipelines across cloud and big data platforms. Develop and execute data validation test cases for batch and incremental data loads. Build and maintain Python-based automation frameworks for data validation. Validate data movement across AWS S3, Glue, Redshift, Lambda , and Databricks environments. Perform source-to-target data validation and data reconciliation. Write complex SQL and PySpark scripts for validating large datasets. Ensure data completeness, consistency, integrity, and accuracy across systems. Identify, troubleshoot, and resolve data quality issues. Collaborate with Data Engineers, Developers, Business Analysts, and Product Owners to understand business requirements. Participate in sprint planning, defect triage, and release validation. Prepare test execution reports, defect reports, and validation documentation. Preferred Qualifications Bachelor’s degree in Computer Science, Information Technology, or a related field. AWS or Databricks certifications are an added advantage. Experience with CI/CD pipelines and version control tools such as Git is preferred. Exposure to orchestration tools like Airflow is a plus. Key Competencies Data Validation ETL Testing Python Automation AWS (S3, Glue, Redshift, Lambda) Databricks PySpark SQL Data Warehousing ETL Validation Frameworks Data Quality & Reconciliation Agile Methodology
More at Qentelli