Padmi
CIM Group logo
CIM Group

real estate development · infrastructure investment

Sr. Machine Learning Ops Engineer

Los Angeles · OnsitePosted 5 months ago
Machine learningSeniorFull Time
Apply at CIM Group

Opens the source posting on jobs.lever.co

Source description

About the role

View original

ML Model Deployment & Platform Management

Lead the design, implementation, and ongoing maintenance of scalable ML infrastructure on Databricks, including ML flow for experiment tracking, model registry, and model serving endpoints.

Oversee the development of the ML Ops platform and automated pipelines for deploying, monitoring, and maintaining models within production environments.

Implement robust solutions for model versioning, systematic retraining, and comprehensive artifact management using Databricks Unity Catalog for ML governance.

Design and manage Databricks Feature Store for consistent feature engineering across training and inference pipelines.

Generative AI & LLM Operations

Architect and implement Retrieval-Augmented Generation (RAG) systems for document Q&A, enabling business teams to query fund documents, investor letters, and market research.

Design, deploy, and manage vector database solutions (Databricks Vector Search, Pinecone, or similar) for semantic search and retrieval across enterprise documents.

Lead LLM fine-tuning and customization initiatives, training models like Claude or open-source alternatives with CIM proprietary data while ensuring data privacy and compliance.

Develop and optimize document processing pipelines including PDF parsing, chunking strategies, and embedding generation for RAG applications.

Implement prompt engineering best practices and LLM evaluation frameworks to ensure output quality, relevance, and factual accuracy.

Build guardrails and safety measures for GenAI applications, including hallucination detection, output validation, and source attribution.

Automation & CI/CD Pipelines

Design and implement extensive automation across the ML workflow, covering model training, testing, validation, and deployment using Databricks Workflows and Asset Bundles.

Set up robust CI/CD pipelines for both traditional ML models and GenAI applications, leveraging GitHub Actions, Azure DevOps, or similar tools.

Automate complex data and model workflows utilizing orchestration tools such as Airflow, Prefect, or Databricks Workflows.

Monitoring, Performance & Reliability

Implement comprehensive monitoring and alerting systems for real-time tracking of model performance, data quality, and GenAI output quality.

Utilize specialized tools (Evidently AI, WhyLabs, Prometheus/Grafana) to proactively detect model drift, data quality anomalies, and RAG retrieval degradation.

Develop evaluation frameworks for GenAI applications including relevance scoring, faithfulness metrics, and human feedback loops.

Troubleshoot issues within production environments, including debugging model deployment failures, RAG retrieval issues, and LLM response quality problems.

Data & Feature Engineering Support

Build and maintain sophisticated feature stores on Databricks, ensuring precise alignment between training and inference data pipelines.

Collaborate with data engineers and information architects to build robust ETL pipelines that feed into the Databricks Lakehouse.

Design embedding pipelines and vector index management strategies for RAG applications, including incremental updates and versioning.

Security, Compliance & Trustworthy AI

Integrate robust security measures directly into ML Ops and GenAI pipelines, including access controls via Unity Catalog and data encryption.

Implement Trustworthy AI guardrails addressing bias detection, explainability, prompt injection prevention, and responsible AI practices.

Ensure GenAI applications handling sensitive fund and investor data comply with regulatory requirements and internal policies.

Collaborate with Legal and Compliance to establish AI governance policies and audit trails for model decisions.

Collaboration & Business Partnership

Engage in extensive collaboration with data scientists, platform engineers, information architects, and DevOps teams to ensure seamless ML/AI integration.

Partner with business teams (Fund Accounting, FP&A, Investor Relations, Sales, Investments) to identify high-value AI use cases and translate business needs into technical solutions.

Communicate complex AI concepts in business terms, managing expectations and demonstrating ROI of ML/GenAI initiatives.

Provide technical mentorship to team members, including refactoring data scientist code for production readiness.

More at CIM Group

Related open roles

View all roles