Source description
About the role
As a Databricks Data Engineer with strong Life Sciences domain experience, you will be responsible for designing, developing, and deploying end-to-end data pipelines and data products using Databricks on AWS. Your role will involve building and optimizing ETL/ELT workflows using Python, PySpark, and SQL for large-scale structured and unstructured datasets. Additionally, you will work extensively with Delta Lake architecture to implement efficient data storage, versioning, and performance optimization. Key Responsibilities: - Design, develop, and deploy end-to-end data pipelines and data products using Databricks on AWS - Build and optimize ETL/ELT workflows using Python, PySpark, and SQL for large-scale structured and unstructured datasets - Work extensively with Delta Lake architecture, implementing efficient data storage, versioning, and performance optimization - Develop and manage Databricks Workflows, Jobs, and Delta Live Tables (DLT) for reliable and scalable pipeline orchestration - Configure and manage Databricks environments (clusters, autoscaling, DBFS, notebooks, Unity Catalog) - Integrate data from multiple sources using Kafka, Airflow, and cloud-native ingestion frameworks - Collaborate with business stakeholders and clients to translate requirements into scalable data solutions - Ensure data quality, governance, and security, especially in regulated Life Sciences environments - Continuously improve existing pipelines through performance tuning, cost optimization, and reliability enhancements Qualification Required: - Strong hands-on experience with: - Python / PySpark - SQL (advanced querying and optimization) - Databricks (Workspace, Notebooks, Clusters, Autoscaling, DBFS) - Deep expertise in: - Delta Lake - Databricks Workflows, Jobs, and Delta Live Tables (DLT) - Unity Catalog (data governance and access control) - Experience working on AWS cloud platform - Hands-on experience with Kafka and Apache Airflow for data ingestion and orchestration - Proven ability to build data pipelines and data products from scratch - Experience working in client-facing roles, with strong communication and requirement-gathering skills - Prior experience in Life Sciences / Pharma domain Additional Details: The company values innovation and collaboration in a high-impact, production environment. If you are passionate about leveraging your technical expertise in Databricks and AWS to create robust data solutions while engaging with stakeholders and understanding business requirements, this is the perfect opportunity for you. As a Databricks Data Engineer with strong Life Sciences domain experience, you will be responsible for designing, developing, and deploying end-to-end data pipelines and data products using Databricks on AWS. Your role will involve building and optimizing ETL/ELT workflows using Python, PySpark, and SQL for large-scale structured and unstructured datasets. Additionally, you will work extensively with Delta Lake architecture to implement efficient data storage, versioning, and performance optimization. Key Responsibilities: - Design, develop, and deploy end-to-end data pipelines and data products using Databricks on AWS - Build and optimize ETL/ELT workflows using Python, PySpark, and SQL for large-scale structured and unstructured datasets - Work extensively with Delta Lake architecture, implementing efficient data storage, versioning, and performance optimization - Develop and manage Databricks Workflows, Jobs, and Delta Live Tables (DLT) for reliable and scalable pipeline orchestration - Configure and manage Databricks environments (clusters, autoscaling, DBFS, notebooks, Unity Catalog) - Integrate data from multiple sources using Kafka, Airflow, and cloud-native ingestion frameworks - Collaborate with business stakeholders and clients to translate requirements into scalable data solutions - Ensure data quality, governance, and security, especially in regulated Life Sciences environments - Continuously improve existing pipelines through performance tuning, cost optimization, and reliability enhancements Qualification Required: - Strong hands-on experience with: - Python / PySpark - SQL (advanced querying and optimization) - Databricks (Workspace, Notebooks, Clusters, Autoscaling, DBFS) - Deep expertise in: - Delta Lake - Databricks Workflows, Jobs, and Delta Live Tables (DLT) - Unity Catalog (data governance and access control) - Experience working on AWS cloud platform - Hands-on experience with Kafka and Apache Airflow for data ingestion and orchestration - Proven ability to build data pipelines and data products from scratch - Experience working in client-facing roles, with strong communication and requirement-gathering skills - Prior experience in Life Sciences / Pharma domain Additional Details: The company values innovation and collaboration in a high-impact, production environment. If you are passionate about leveraging your technical expertise in Databr
More at GoStravvy