Source description
About the role
Business Partnership & Requirements Gathering
Partners extensively with business stakeholders across Fund Accounting, FP&A, Global Client Group, and Investments teams to understand data requirements, pain points, and use cases.
Translate complex business requirements into robust data models, asking questions first before proposing solutions to ensure alignment with actual business needs.
Collaborate with data analysts, data scientists, MLOps engineers, and application developers to understand technical requirements and ensure models support downstream use cases.
Build trust and influence across teams by demonstrating business value and explaining technical constraints in accessible terms.
Data Model Design & Development
Design and develop conceptual, logical, and physical data models for various data initiatives, including data warehouses, data lakes, operational data stores, and transactional systems.
Specialize in designing highly optimized dimensional models (star schemas, snowflake schemas) for analytical reporting and business intelligence applications.
Apply various data modeling techniques as appropriate: Dimensional Modeling (Kimball), 3NF (Inmon), Data Vault, and NoSQL modeling patterns.
Databricks Lakehouse Architecture
Design and implement medallion architecture (bronze/silver/gold) patterns within the Databricks Lakehouse, establishing standards where none currently exist.
Optimize data models leveraging Delta Lake features including ACID transactions, time travel, schema evolution, Z-ordering, and liquid clustering.
Design partitioning strategies that balance query performance with file management, avoiding over-partitioning while enabling partition pruning.
Implement Unity Catalog namespace hierarchy (catalog, schema, table) for multi-domain, multi-environment data organization and governance.
Collaborate with MLOps engineers on data models that support ML feature stores and GenAI/RAG applications.
Multi-Platform Data Architecture
Design and maintain relational database schemas for both operational and analytical workloads within Snowflake and other RDBMS, ensuring integration with Databricks Lakehouse.
Design and optimize data structures for NoSQL databases (e.g., MongoDB), considering document structures, indexing, and query patterns for specific application needs.
Develop frameworks for deciding when data belongs in Databricks Lakehouse vs. Snowflake vs. MongoDB based on workload characteristics and use cases.
Performance Optimization
Provide input and recommendations on query optimization, indexing strategies, and data partitioning based on data model design.
Diagnose and resolve performance issues including data skew, small files problems, and inefficient join strategies in Spark/Databricks environments.
Collaborate with database administrators and data engineers on OPTIMIZE, VACUUM, and ANALYZE strategies for Delta tables.
ETL/ELT Collaboration & Data Pipeline Design
Work closely with Data Engineers to ensure data models are efficiently implemented and align with ETL/ELT processes using Auto Loader, Delta Live Tables, or traditional Spark jobs.
Provide guidance on data mapping, transformation rules, schema evolution handling, and data loading strategies.
Design slowly changing dimension (SCD) patterns using Delta Lake MERGE operations and Change Data Feed for downstream propagation.
Data Governance & Standards
Establish and enforce data modeling standards, naming conventions, metadata management, and data governance policies—building these foundations where they do not currently exist.
Implement row-level and column-level security patterns using Unity Catalog for sensitive fund and investor data.
Design and maintain data lineage tracking from source systems through bronze/silver/gold layers to final reports.
Contribute to the development and maintenance of a comprehensive data dictionary and metadata repository.
Operate within compliance framework to ensure ethical data handling, regulatory compliance, and consistency across the enterprise.
Documentation & Knowledge Management
Create and maintain detailed data model documentation, including data dictionaries, entity-relationship diagrams (ERDs), data flow diagrams, and data lineage.
Document Lakehouse design patterns, medallion architecture implementations, and platform-specific best practices for team knowledge sharing.
Contribute to data quality initiatives by identifying potential data quality issues at the modeling stage and collaborating on solutions.
More at CIM Group
