Source description
About the role
Role Summary Data Engineer Job Description The Data Engineer is responsible for designing, building, and maintaining scalable data pipelines, data platforms, and integration solutions across cloud environments. This role focuses on transforming raw data into reliable, high-quality datasets to support analytics, reporting, and AI/ML use cases. Key Responsibilities Data Engineering & Pipeline Development Design, develop, and maintain ETL/ELT pipelines for data ingestion, transformation, and loadingBuild scalable data workflows using batch and real-time processing frameworksDevelop and optimize data pipelines for performance, reliability, and scalabilityHandle structured and unstructured data across multiple sources Data Platform & Cloud Implementation Work with cloud platforms (Azure / AWS / GCP) to build and manage data solutionsUtilize cloud-native services such as:Data lakes, warehouses, and lakehouse platformsDistributed compute (e.g., Spark, Databricks, Synapse)Support deployment and management of data infrastructure and storage systems Data Integration & Transformation Integrate data from multiple systems including APIs, databases, applications, and streaming sourcesImplement transformation logic using SQL, PySpark, or other data processing toolsEnsure consistency and accuracy across data pipelines Data Quality, Governance & Security Implement data validation, cleansing, and quality checksEnsure compliance with data governance, privacy, and security policies (PII/PHI handling)Maintain data lineage, metadata, and documentation Monitoring, Optimization & Reliability Monitor data pipelines and workflows for failures and performance issuesImplement logging, alerting, and troubleshooting mechanismsOptimize pipelines for cost, speed, and resource utilization Collaboration & Support Work closely with data architects, analysts, and business stakeholders to understand requirementsSupport analytics, BI, and AI teams with clean and reliable datasetsParticipate in code reviews, testing, and deployment processes Documentation & Best Practices Document data flows, pipeline logic, and technical designsFollow best practices for data modeling, schema design, and version controlMaintain reusable components and frameworks Required Experience 38 years of experience in Data Engineering or related rolesStrong experience with:ETL/ELT tools and frameworksSQL and data modeling conceptsPython / PySpark / Scala (at least one)Hands-on experience with:Cloud platforms (Azure / AWS / GCP)Big data tools (Spark, Databricks, Synapse, etc.)Experience with data streaming tools (Kafka/Event Hubs) is a plusUnderstanding of CI/CD and DevOps practices Role Summary Data Engineer Job Description The Data Engineer is responsible for designing, building, and maintaining scalable data pipelines, data platforms, and integration solutions across cloud environments. This role focuses on transforming raw data into reliable, high-quality datasets to support analytics, reporting, and AI/ML use cases. Key Responsibilities Data Engineering & Pipeline Development Design, develop, and maintain ETL/ELT pipelines for data ingestion, transformation, and loadingBuild scalable data workflows using batch and real-time processing frameworksDevelop and optimize data pipelines for performance, reliability, and scalabilityHandle structured and unstructured data across multiple sources Data Platform & Cloud Implementation Work with cloud platforms (Azure / AWS / GCP) to build and manage data solutionsUtilize cloud-native services such as:Data lakes, warehouses, and lakehouse platformsDistributed compute (e.g., Spark, Databricks, Synapse)Support deployment and management of data infrastructure and storage systems Data Integration & Transformation Integrate data from multiple systems including APIs, databases, applications, and streaming sourcesImplement transformation logic using SQL, PySpark, or other data processing toolsEnsure consistency and accuracy across data pipelines Data Quality, Governance & Security Implement data validation, cleansing, and quality checksEnsure compliance with data governance, privacy, and security policies (PII/PHI handling)Maintain data lineage, metadata, and documentation Monitoring, Optimization & Reliability Monitor data pipelines and workflows for failures and performance issuesImplement logging, alerting, and troubleshooting mechanismsOptimize pipelines for cost, speed, and resource utilization Collaboration & Support Work closely with data architects, analysts, and business stakeholders to understand requirementsSupport analytics, BI, and AI teams with clean and reliable datasetsParticipate in code reviews, testing, and deployment processes Documentation & Best Practices Document data flows, pipeline logic, and technical designsFollow best practices for data modeling, schema design, and version controlMaintain reusable components and frameworks Required Experience 38 years of experience in D
More at Anblicks