Source description
About the role
As a Data Engineer, you will be responsible for designing and implementing Medallion architecture with a focus on schema enforcement, audit trails, versioning, time travel, and incremental processing strategies. You will also design and maintain optimized data models for Lakehouse and warehouse consumption. Additionally, your key responsibilities will include: - Developing high-performance distributed data pipelines using Apache Spark - Optimizing Spark workloads through partitioning, caching, broadcast joins, and query tuning - Implementing efficient incremental data processing and change data capture strategies - Monitoring and troubleshooting Spark job failures, latency, and resource bottlenecks - Designing and implementing CI/CD pipelines for data engineering workflows - Automating build, test, and deployment of data pipelines across environments Qualifications Required: - 5+ years of experience building distributed data pipelines - Strong expertise in Python and SQL - Experience with Apache Spark (PySpark/Scala Spark) and Spark performance tuning - Knowledge of partitioning strategies, file optimization, execution plan analysis, and query optimization - Familiarity with Delta Lake / Lakehouse architectures, incremental processing, and time travel - Proficiency in CI/CD tools (Azure DevOps / GitHub Actions / Jenkins or similar) and infrastructure-as-code concepts - Experience with Azure/AWS/GCP data platforms and workflow orchestration tools (Airflow or similar) Additional Details: - Good to have experience in Microsoft Fabric / Databricks, streaming frameworks, data observability tools, and working in the insurance or financial domain. As a Data Engineer, you will be responsible for designing and implementing Medallion architecture with a focus on schema enforcement, audit trails, versioning, time travel, and incremental processing strategies. You will also design and maintain optimized data models for Lakehouse and warehouse consumption. Additionally, your key responsibilities will include: - Developing high-performance distributed data pipelines using Apache Spark - Optimizing Spark workloads through partitioning, caching, broadcast joins, and query tuning - Implementing efficient incremental data processing and change data capture strategies - Monitoring and troubleshooting Spark job failures, latency, and resource bottlenecks - Designing and implementing CI/CD pipelines for data engineering workflows - Automating build, test, and deployment of data pipelines across environments Qualifications Required: - 5+ years of experience building distributed data pipelines - Strong expertise in Python and SQL - Experience with Apache Spark (PySpark/Scala Spark) and Spark performance tuning - Knowledge of partitioning strategies, file optimization, execution plan analysis, and query optimization - Familiarity with Delta Lake / Lakehouse architectures, incremental processing, and time travel - Proficiency in CI/CD tools (Azure DevOps / GitHub Actions / Jenkins or similar) and infrastructure-as-code concepts - Experience with Azure/AWS/GCP data platforms and workflow orchestration tools (Airflow or similar) Additional Details: - Good to have experience in Microsoft Fabric / Databricks, streaming frameworks, data observability tools, and working in the insurance or financial domain.
More at Talent Toppers
Related open roles
Architect-Python FullStack (React)
Delhi NCR
FullStack Developer(Python)
Delhi NCR
FullStack Developer(Python) (Gurugram)
India
Oracle Fusion HCM Techno-Functional Compensation Sr. Analyst
Delhi NCR
Oracle Fusion ERP Techno-Functional Sr. Analyst
Delhi NCR
Artificial Intelligence Specialist - Machine Learning/Deep Learning
Delhi NCR