Source description
About the role
Role & responsibilities - Provide technical recommendations and evaluate project feasibility and workloads. - Design and implement efficient pipeline architectures for data ingestion, transformation, and storage. - Develop modular, efficient code for data transformation and integration, ensuring adherence to contemporary development practices. - Perform data ingestion from source systems, including ETL processes, when tables are not available on the data platform. - Conduct regular pipeline maintenance and upgrades, troubleshooting and resolving issues to ensure optimal performance. - Implement data quality testing (automatic) and validation to maintain high standards of data integrity and reliability Preferred candidate profile - Expert knowledge in dbt, SQL, Python, Spark (PySpark), Git, Azure DevOps and Databricks. - Expertise in cloud engineering (e.g., AWS, Azure); experience with infrastructureascode is a plus, and deploying CI/CD pipelines is a must. - Strong problemsolving skills. - 5 years of experience in creating and maintaining data pipelines. - Experience handling scientific datasets for cheminformatics or bioinformatics (e.g., microbiome) is a plus. - Experience with FAIR data principles, data governance, and data standards commonly used in scientific domains. .
More at DSM Firmenich