Padmi
Xebia logo
Xebia

AI and Machine Learning · Cloud platforms (AWS, Azure)

Lead Scientific Data Systems Engineer

IndiaPosted 2 months ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at Xebia

Opens the source posting on shine.com

Source description

About the role

View original

As a highly specialized Principal Scientific Data Architect within the Google Cloud Platform (GCP) ecosystem, your role is crucial in bridging the gap between advanced Google Cloud engineering and life sciences discovery. You will play a key role in redefining how scientific data is structured, scaled, and consumed across various divisions, enabling the design of data systems that directly power in-silico molecular discovery and autonomous Agentic AI frameworks. Key Responsibilities: - GCP-Native Data Architecture & Paradigm Shifts - Design and implement version-controlled, programmatically managed data schemas integrated with Google BigQuery. - Treat data assets with software engineering rigor by implementing data versioning, programmability, and automated quality testing. - Architect highly optimized, metadata-driven, configuration-led data pipelines using Google Cloud Composer or Dataflow. - Scientific Domain Integration - Translate complex biological and chemical concepts into highly scalable logical and physical data models within BigQuery and Databricks. - Collaborate with computational chemists, biologists, and AI engineers to support predictive in-silico modeling. - Design robust data layouts for autonomous AI agents to extract properties and explain molecular behavior. - Platform & Ecosystem Strategy - Optimize interoperability between Databricks on GCP and Google BigQuery storage and analytics. - Integrate semantic web technologies and knowledge graphs into the Google Cloud data fabric. - Ensure data availability and high-performance querying for downstream multi-agent AI ecosystems. Required Skills & Qualifications: - Scientific Domain Knowledge - Robust background or proven experience working inside life sciences, pharmaceuticals, biotech, or scientific research organizations. - Ability to converse fluently with scientists regarding therapeutic modalities, molecular properties, and R&D pipelines. - GCP & Technical Architecture Expertise - Mastery of Google BigQuery and Databricks on GCP. - Proven track record of implementing Schema as Code and Data as Code paradigms. - Deep experience with configuration-driven pipeline orchestrators, specifically Google Cloud Composer / Apache Airflow. - Strong understanding of relational, dimensional, and graph-based data modeling. - Soft Skills & Leadership - Ability to conceptualize complex in-silico data solutions at a high strategic level. - Exceptional communication skills to articulate the business and scientific value of data architecture. Preferred Qualifications: - Professional Google Cloud Data Engineer or Google Cloud Professional Cloud Architect certification. - Degree in Computer Science, Data Engineering, Bioinformatics, Computational Chemistry, or a related quantitative field. - Experience setting up GCP data foundations engineered for Large Language Models and autonomous AI agents. As a highly specialized Principal Scientific Data Architect within the Google Cloud Platform (GCP) ecosystem, your role is crucial in bridging the gap between advanced Google Cloud engineering and life sciences discovery. You will play a key role in redefining how scientific data is structured, scaled, and consumed across various divisions, enabling the design of data systems that directly power in-silico molecular discovery and autonomous Agentic AI frameworks. Key Responsibilities: - GCP-Native Data Architecture & Paradigm Shifts - Design and implement version-controlled, programmatically managed data schemas integrated with Google BigQuery. - Treat data assets with software engineering rigor by implementing data versioning, programmability, and automated quality testing. - Architect highly optimized, metadata-driven, configuration-led data pipelines using Google Cloud Composer or Dataflow. - Scientific Domain Integration - Translate complex biological and chemical concepts into highly scalable logical and physical data models within BigQuery and Databricks. - Collaborate with computational chemists, biologists, and AI engineers to support predictive in-silico modeling. - Design robust data layouts for autonomous AI agents to extract properties and explain molecular behavior. - Platform & Ecosystem Strategy - Optimize interoperability between Databricks on GCP and Google BigQuery storage and analytics. - Integrate semantic web technologies and knowledge graphs into the Google Cloud data fabric. - Ensure data availability and high-performance querying for downstream multi-agent AI ecosystems. Required Skills & Qualifications: - Scientific Domain Knowledge - Robust background or proven experience working inside life sciences, pharmaceuticals, biotech, or scientific research organizations. - Ability to converse fluently with scientists regarding therapeutic modalities, molecular properties, and R&D pipelines. - **GCP & Technical Architect

One address, no account. We’ll tell you when matching roles go live.

More at Xebia

Related open roles

View all roles