Source description
About the role
About the role: Engineer - Data & AI We are looking for a skilled and motivated Data Engineer with 4+ years of experience in designing, developing, and maintaining scalable data integration and analytics solutions. The ideal candidate will have strong expertise in Qlik Replicate, GCP Data Engineering services, PySpark, SQL, and Apache Airflow, along with hands-on experience in building CDC and ETL pipelines within cloud-based data platforms. This role involves working on modern data architectures, enabling seamless data movement from enterprise source systems to cloud environments, and supporting robust analytics and reporting capabilities. Key Responsibilities Design, develop, and maintain Change Data Capture (CDC) pipelines using Qlik Replicate for real-time and batch data integration. Implement and support data replication from IBM DB2 and other enterprise data sources to Google BigQuery, Google Cloud Storage (GCS), and PostgreSQL. Build scalable and optimized data transformation pipelines using PySpark and SQL. Develop and manage workflow orchestration and scheduling using Apache Airflow. Design and implement data solutions following the Medallion Architecture (Bronze, Silver, Gold) framework. Leverage Google Cloud Platform (GCP) services to build high-performance and reliable data engineering solutions. Monitor, troubleshoot, and optimize data pipelines to ensure high availability and data quality. Collaborate with business stakeholders, analysts, and cross-functional teams to translate data requirements into scalable technical solutions. Implement best practices for data governance, performance optimization, and operational excellence. Support production deployments, incident management, and continuous improvement initiatives. Roles and Responsibilities Required Skills & Experience Technical Skills Strong hands-on experience with Qlik Replicate for CDC and data replication. Experience working with Google Cloud Platform (GCP) data services. Proficiency in PySpark for large-scale data processing. Advanced knowledge of SQL and query optimization techniques. Experience with Apache Airflow for workflow orchestration and monitoring. Hands-on experience with BigQuery and Google Cloud Storage (GCS). Strong understanding of relational databases including IBM DB2 and PostgreSQL. Experience designing and implementing ETL/ELT pipelines. Good understanding of Data Warehousing concepts and Medallion Architecture. Preferred Skills Experience with cloud-native data lake and analytics solutions. Familiarity with CI/CD practices and DevOps methodologies. Knowledge of data quality, governance, and performance tuning best practices. Exposure to Agile development methodologies. Qualifications Bachelor's or master’s degree in computer science, Information Technology, Engineering, or a related field. 4+ years of experience in Data Engineering, Data Integration, or related areas. Proven experience delivering enterprise-scale data migration, replication, and analytics solutions.
More at Exponentia.ai Private Limited