Padmi
Institute of Foundation Models logo
Institute of Foundation Models

foundation models · large language models

Eval360 - Error Analysis Engineer

San Francisco Bay Area · Onsite$150k–$450k/yrPosted 3 months ago
Software engineeringSeniorFull Time
Apply at Institute of Foundation Models

Opens the source posting on jobs.lever.co

Source description

About the role

View original

• Collaborate with researchers, machine learning engineers, data scientists, product managers, and internal stakeholders to implement innovative software solutions for Eval360 and related model evaluation workflows.

• Build and improve Eval360 as an evaluation service that acts as a quality gate for model development, model comparison, and model release decisions.

• Perform deep error analysis on model outputs, including identifying failure patterns, categorizing issues, tracing root causes, and proposing improvements to evaluation methodology.

• Develop tools, workflows, and dashboards that make it easier for researchers and engineers to inspect model failures, compare model behavior, and understand quality regressions.

• Design and implement client-side and server-side architecture for evaluation review systems, error analysis interfaces, reporting tools, and internal evaluation applications.

• Develop responsive, usable interfaces that support error triage, annotation review, evaluation debugging, and model quality investigation.

• Build and maintain back-end services, APIs, data pipelines, and integrations that support evaluation execution, results storage, analysis, and reporting.

• Test software to ensure responsiveness, correctness, reliability, and efficiency across evaluation workflows.

• Troubleshoot, debug, and upgrade evaluation systems, including identifying issues in data processing, evaluation metrics, model output handling, job orchestration, and user-facing analysis tools.

• Create and maintain security, access control, and data protection settings for evaluation data, model outputs, annotations, and internal tooling.

• Write clear technical documentation for Eval360 systems, error taxonomies, evaluation workflows, debugging procedures, and user-facing tools.

• Work with researchers, data scientists, analysts, and machine learning engineers to improve evaluation quality, model diagnostics, and failure-mode visibility.

• Keep track of new development tools, evaluation frameworks, model analysis methods, data quality techniques, and architectures relevant to AI evaluation systems.

• Contribute to the design of error taxonomies, evaluation rubrics, quality thresholds, regression detection methods, and model readiness criteria.

• Help ensure Eval360 produces reliable, interpretable, and actionable signals for model quality gates.

• Contribute to research publications, technical reports, internal knowledge sharing, and external presentations where appropriate.

• Contribute to intellectual property and thought leadership in AI evaluation, error analysis, model quality measurement, and evaluation infrastructure.

• Perform all other duties as reasonably directed by the line manager that are aligned with these functional objectives.

One address, no account. We’ll tell you when matching roles go live.

More at Institute of Foundation Models

Related open roles

View all roles