Padmi

Python Developers

ChennaiPosted 2 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Bahwan CyberTek

Opens the source posting on shine.com

Source description

About the role

View original

Strong hands-on experience with Python and PySpark Solid SQL skills for complex transformations, validation, and reconciliation Experience with OCR tools/libraries for extracting tables from scanned or unstructured PDFs Experience building ETL pipelines from API and web-based sources Working knowledge of Medallion Architecture / layered data lake design Hands-on experience with Apache Airflow (workflow orchestration) Hands-on experience with Apache NiFi (data flow automation) Experience with Git and CI/CD pipeline setup Strong understanding of data validation, quality checks, and reconciliation practices Design and develop ETL/ELT pipelines using Python and PySpark to ingest data from APIs, websites, emails,and internal systems into a centralized data lake Build and maintain OCR-based extraction pipelines to pull tabular data from large, unstructured PDF documents Implement and manage a Medallion Architecture (Bronze/Silver/Gold layers) for staged data refinement raw ingestion, validation/cleansing, and business-ready aggregation Write complex SQL for data validation, reconciliation, and querying across large datasets Orchestrate and schedule pipeline workflows using Apache Airflow Design data flow automation and routing using Apache NiFi Build and maintain CI/CD pipelines (Git-based) for version-controlled, repeatable deployments Ensure data accuracy, completeness, and timeliness against fixed delivery/reporting schedules Collaborate with cross-functional teams to understand data requirements and troubleshoot pipeline issues Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying. Strong hands-on experience with Python and PySpark Solid SQL skills for complex transformations, validation, and reconciliation Experience with OCR tools/libraries for extracting tables from scanned or unstructured PDFs Experience building ETL pipelines from API and web-based sources Working knowledge of Medallion Architecture / layered data lake design Hands-on experience with Apache Airflow (workflow orchestration) Hands-on experience with Apache NiFi (data flow automation) Experience with Git and CI/CD pipeline setup Strong understanding of data validation, quality checks, and reconciliation practices Design and develop ETL/ELT pipelines using Python and PySpark to ingest data from APIs, websites, emails,and internal systems into a centralized data lake Build and maintain OCR-based extraction pipelines to pull tabular data from large, unstructured PDF documents Implement and manage a Medallion Architecture (Bronze/Silver/Gold layers) for staged data refinement raw ingestion, validation/cleansing, and business-ready aggregation Write complex SQL for data validation, reconciliation, and querying across large datasets Orchestrate and schedule pipeline workflows using Apache Airflow Design data flow automation and routing using Apache NiFi Build and maintain CI/CD pipelines (Git-based) for version-controlled, repeatable deployments Ensure data accuracy, completeness, and timeliness against fixed delivery/reporting schedules Collaborate with cross-functional teams to understand data requirements and troubleshoot pipeline issues Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

One address, no account. We’ll tell you when matching roles go live.

More at Bahwan CyberTek

Related open roles

View all roles