Source description
About the role
Role Overview: As a Senior Data Engineer, your main responsibility will be designing, building, and optimizing data pipelines and lakehouse architectures on AWS. You will play a crucial role in ensuring data availability, quality, lineage, and governance across analytical and operational platforms. Your expertise will be instrumental in developing scalable, secure, and cost-effective data solutions that drive advanced analytics and business intelligence. Key Responsibilities: - Implement and manage S3 (raw, staging, curated zones), Glue Catalog, Lake Formation, and Iceberg/Hudi/Delta Lake for schema evolution and versioning. - Develop PySpark jobs on Glue/EMR, enforce schema validation, partitioning, and scalable transformations. - Build workflows using Step Functions, EventBridge, or Airflow (MWAA), with CI/CD deployments via CodePipeline & CodeBuild. - Apply schema contracts, validations (Glue Schema Registry, Deequ, Great Expectations), and maintain lineage/metadata using Glue Catalog or third-party tools (Atlan, OpenMetadata, Collibra). - Enable Athena and Redshift Spectrum queries, manage operational stores (DynamoDB/Aurora), and integrate with OpenSearch for observability. - Design efficient partitioning/bucketing strategies, adopt columnar formats (Parquet/ORC), and implement spot instance usage/bookmarking. - Enforce IAM-based access policies, apply KMS encryption, private endpoints, and GDPR/PII data masking. - Prepare Gold-layer KPIs for dashboards, forecasting, and customer insights with QuickSight, Superset, or Metabase. - Partner with analysts, data scientists, and DevOps to enable seamless data consumption and delivery. Qualifications Required: - Hands-on expertise with AWS data stack (S3, Glue, Lake Formation, Athena, Redshift, EMR, Lambda). - Strong programming skills in PySpark & Python for ETL, scripting, and automation. - Proficiency in SQL (CTEs, window functions, complex aggregations). - Experience in data governance, quality frameworks (Deequ, Great Expectations). - Knowledge of data modeling, partitioning strategies, and schema enforcement. - Familiarity with BI integration (QuickSight, Superset, Metabase). - Bachelor's/Master's degree in Computer Science, Information Technology, or related field. - Minimum 4 years of proven experience in data engineering with AWS. (Note: The additional details of the company were not included in the provided Job Description.) Role Overview: As a Senior Data Engineer, your main responsibility will be designing, building, and optimizing data pipelines and lakehouse architectures on AWS. You will play a crucial role in ensuring data availability, quality, lineage, and governance across analytical and operational platforms. Your expertise will be instrumental in developing scalable, secure, and cost-effective data solutions that drive advanced analytics and business intelligence. Key Responsibilities: - Implement and manage S3 (raw, staging, curated zones), Glue Catalog, Lake Formation, and Iceberg/Hudi/Delta Lake for schema evolution and versioning. - Develop PySpark jobs on Glue/EMR, enforce schema validation, partitioning, and scalable transformations. - Build workflows using Step Functions, EventBridge, or Airflow (MWAA), with CI/CD deployments via CodePipeline & CodeBuild. - Apply schema contracts, validations (Glue Schema Registry, Deequ, Great Expectations), and maintain lineage/metadata using Glue Catalog or third-party tools (Atlan, OpenMetadata, Collibra). - Enable Athena and Redshift Spectrum queries, manage operational stores (DynamoDB/Aurora), and integrate with OpenSearch for observability. - Design efficient partitioning/bucketing strategies, adopt columnar formats (Parquet/ORC), and implement spot instance usage/bookmarking. - Enforce IAM-based access policies, apply KMS encryption, private endpoints, and GDPR/PII data masking. - Prepare Gold-layer KPIs for dashboards, forecasting, and customer insights with QuickSight, Superset, or Metabase. - Partner with analysts, data scientists, and DevOps to enable seamless data consumption and delivery. Qualifications Required: - Hands-on expertise with AWS data stack (S3, Glue, Lake Formation, Athena, Redshift, EMR, Lambda). - Strong programming skills in PySpark & Python for ETL, scripting, and automation. - Proficiency in SQL (CTEs, window functions, complex aggregations). - Experience in data governance, quality frameworks (Deequ, Great Expectations). - Knowledge of data modeling, partitioning strategies, and schema enforcement. - Familiarity with BI integration (QuickSight, Superset, Metabase). - Bachelor's/Master's degree in Computer Science, Information Technology, or related field. - Minimum 4 years of proven experience in data engineering with AWS. (Note: The additional details of the company were not included in the provided Job Description.)
More at LeewayHertz