Padmi
Institute of Foundation Models logo
Institute of Foundation Models

foundation models · large language models

Eval360 - Error Analysis Engineer

San Francisco Bay Area · Onsite$150k–$450k/yrPosted 1 month ago
Machine learningUnspecifiedFull Time
Apply at Institute of Foundation Models

Opens the source posting on jobs.lever.co

Source description

About the role

View original

• Collaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows.

• Build and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions.

• Perform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology.

• Develop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions.

• Design and implement client-side and server-side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications.

• Develop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation.

• Build and maintain back-end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting.

• Test software to ensure responsiveness, correctness, reliability, and efficiency across evaluation workflows.

• Troubleshoot, debug, and upgrade evaluation systems, including identifying issues in data processing, evaluation metrics, model output handling, job orchestration, and user-facing analysis tools.

• Create and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling.

• Write clear technical documentation for Eval360 systems, error taxonomies, evaluation workflows, debugging procedures, and user-facing tools.

• Work with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure-mode visibility.

• Keep track of new development tools, evaluation frameworks, model analysis methods, data quality techniques, and architectures relevant to AI evaluation systems.

• Contribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria.

• Help ensure Eval360 produces reliable, interpretable, and actionable signals for model quality gates.

• Contribute to research publications, technical reports, internal knowledge sharing, and external presentations where appropriate.

• Contribute to intellectual property and thought leadership in AI evaluation, error analysis, model quality measurement, and evaluation infrastructure.

• Perform all other duties as reasonably directed by the line manager that are aligned with these functional objectives.

More at Institute of Foundation Models

Related open roles

View all roles