Source description
About the role
Data Engineer: key skill - Apache Iceberg, ingestion, data quality Data Engineer – Apache Iceberg, Data Ingestion & Data Quality Job Title Data Engineer Experience 8+ Years Job Summary We are seeking an experienced Data Engineer with strong expertise in Apache Iceberg, data ingestion frameworks, and data quality management. The ideal candidate will be responsible for designing, building, and optimizing scalable data pipelines and modern data lake architectures to support enterprise analytics and business intelligence initiatives. The candidate should possess hands-on experience with large-scale data processing, cloud-based data platforms, metadata management, and data governance practices. This role requires close collaboration with data architects, analysts, and business stakeholders to deliver reliable and high-quality data solutions. Key Responsibilities Design, develop, and maintain scalable data ingestion pipelines for batch and real-time data processing.
Implement and manage data lakehouse architectures using Apache Iceberg.
Build robust ETL/ELT frameworks to ingest data from multiple structured and unstructured data sources.
Ensure data quality, consistency, accuracy, and completeness across enterprise data platforms.
Develop automated data validation, reconciliation, and monitoring processes.
Optimize data storage, partitioning, compaction, and query performance in Iceberg tables.
Work with large-scale distributed processing frameworks to handle high-volume datasets.
Implement data governance, metadata management, lineage tracking, and auditing mechanisms.
Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
Troubleshoot data pipeline failures and performance bottlenecks.
Support data migration and modernization initiatives involving data warehouses and data lakes.
Implement CI/CD practices for data engineering workflows and deployments.
Create and maintain technical documentation, architecture diagrams, and operational procedures.
Mandatory Skills Apache Iceberg Strong hands-on experience with Apache Iceberg table format.
Expertise in partition evolution, schema evolution, time travel, and snapshot management.
Experience optimizing Iceberg tables for performance and storage efficiency.
Understanding of metadata management and transactional consistency.
Data Ingestion Experience designing and implementing batch and streaming ingestion pipelines.
Expertise in ETL/ELT development and orchestration.
Experience integrating data from databases, APIs, files, cloud storage, and event streams.
Knowledge of CDC (Change Data Capture) concepts and implementation.
Data Quality Strong experience implementing data quality frameworks and controls.
Data validation, profiling, cleansing, reconciliation, and anomaly detection.
Experience defining data quality metrics, SLAs, and monitoring dashboards.
Knowledge of data governance and master data management principles.
Technical Skills Programming: Python, SQL, Scala, or Java
Data Processing: Apache Spark, PySpark
Data Storage: Apache Iceberg, Delta Lake, Parquet, ORC
Workflow Orchestration: Apache Airflow, Prefect, or similar tools
Streaming Technologies: Apache Kafka, Spark Streaming
Cloud Platforms: AWS, Azure, or GCP
Data Warehousing: Snowflake, BigQuery, Redshift, Synapse, or Databricks
Version Control: Git
CI/CD Tools: Jenkins, GitHub Actions, Azure DevOps, or similar
Preferred Qualifications
-
Experience with modern Lakehouse architecture.
-
Exposure to data governance and catalog tools.
-
Knowledge of data security, encryption, and compliance requirements.
-
Experience working in Agile/Scrum environments.
-
Relevant cloud or data engineering certifications are preferred.
-
Soft Skills Strong analytical and problem-solving abilities.
-
Excellent communication and stakeholder management skills.
-
Ability to work independently and collaboratively in a fast-paced environment.
-
Strong ownership mindset and attention to detail.
More at Highbrow Technology