Source description
About the role
Responsibilities: Lead enterprise-scale data engineering initiatives and cloud modernisation programmes.Design scalable batch and real-time data pipelines across AWS and GCP platforms.Drive Data Lake and Lakehouse architecture implementations using Databricks and Snowflake.Mentor engineering teams and establish engineering best practices.Collaborate with architects, DevOps, business stakeholders, and cross-functional teams.Design, develop, and maintain scalable ETL/ELT pipelines.Build batch and real-time data ingestion frameworks.Develop optimised SQL transformations and scalable data processing solutions.Implement distributed data processing using Spark/PySpark.Perform query optimisation, partitioning, clustering, and workload tuning.Implement orchestration, monitoring, alerting, and operational support.Ensure data quality, governance, lineage, and metadata management.Support troubleshooting and production issue resolution.Enable business intelligence and analytics use cases through scalable curated datasets and optimised data models.Collaborate with BI and analytics teams supporting reporting platforms such as Power BI, AWS QuickSight, or equivalent visualisation tools. Requirements: Required Skills: AWS, GCP, Databricks, Snowflake, Spark, Delta Lake, Python, SQL, PySpark, Airflow, Kafka, AWS Glue, Dataflow, Redshift, Athena, Power BI/QuickSight, AWS QuickSight, or equivalent analytics/reporting platforms.Familiar with ETL/ELT frameworks and distributed data processing.Experience in CI/CD, Terraform, Docker, and Kubernetes.Know about data governance, security, monitoring, and cost optimisation. .