Source description
About the role
Design and develop ETL/ELT pipelines using Python and PySpark to ingest data from APIs, websites, emails,and internal systems into a centralized data lake Build and maintain OCR-based extraction pipelines to pull tabular data from large, unstructured PDF documents Implement and manage a Medallion Architecture (Bronze/Silver/Gold layers) for staged data refinement — raw ingestion, validation/cleansing, and business-ready aggregation Write complex SQL for data validation, reconciliation, and querying across large datasets Orchestrate and schedule pipeline workflows using Apache Airflow Design data flow automation and routing using Apache NiFi Build and maintain CI/CD pipelines (Git-based) for version-controlled, repeatable deployments Ensure data accuracy, completeness, and timeliness against fixed delivery/reporting schedules Collaborate with cross-functional teams to understand data requirements and troubleshoot pipeline issues
Strong hands-on experience with Python and PySpark Solid SQL skills for complex transformations, validation, and reconciliation Experience with OCR tools/libraries for extracting tables from scanned or unstructured PDFs Experience building ETL pipelines from API and web-based sources Working knowledge of Medallion Architecture / layered data lake design Hands-on experience with Apache Airflow (workflow orchestration) Hands-on experience with Apache NiFi (data flow automation) Experience with Git and CI/CD pipeline setup Strong understanding of data validation, quality checks, and reconciliation practices
More at Bahwan CyberTek