Padmi

LLM Agent Evaluation Engineer

ChennaiPosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at SIMPLIIGENCE INC

Opens the source posting on shine.com

Source description

About the role

View original

We are looking for an Evaluation Engineer to own and scale the evaluation program that gates our LLM and agentic system releases. Please share an updated profile to kavitha@simpliigence.com (+91) 74839 25904 Mode of Work: Remote Required Skills & Experience Eval Ownership: 6+ years software/ML engineering with demonstrated ownership of an LLM/agent evaluation program that gated real releases. Strong Python (pandas, SQL, pytest). Agent (not just LLM) Evals: Hands-on with trajectory/trace-based evaluation, tool-calling and planning metrics, and pass^k multi-run methodology , output-only eval experience is not enough. Judge Validation: Proven ability to build and statistically validate LLM-as-judge / agent-as-judge evaluators and to articulate why naive judges fail on multi-turn traces and how to mitigate it. Tooling: Practical experience across the current stack , e.g., Braintrust (CI regression), Arize Phoenix (OpenTelemetry-native observability), Promptfoo (OWASP red-team), Galileo, DeepEval/Confident AI, Ragas, Langfuse/LangSmith , and the judgment to know what each metric proves. Trace Fluency: Comfort reading OpenTelemetry-style agent traces (spans for retrieval, tool calls, sub-agent handoffs) and attributing failures to components. Skeptical Communication: Defends uncomfortable numbers to delivery leadership and client stakeholders; influences engineering without authority. We are looking for an Evaluation Engineer to own and scale the evaluation program that gates our LLM and agentic system releases. Please share an updated profile to kavitha@simpliigence.com (+91) 74839 25904 Mode of Work: Remote Required Skills & Experience Eval Ownership: 6+ years software/ML engineering with demonstrated ownership of an LLM/agent evaluation program that gated real releases. Strong Python (pandas, SQL, pytest). Agent (not just LLM) Evals: Hands-on with trajectory/trace-based evaluation, tool-calling and planning metrics, and pass^k multi-run methodology , output-only eval experience is not enough. Judge Validation: Proven ability to build and statistically validate LLM-as-judge / agent-as-judge evaluators and to articulate why naive judges fail on multi-turn traces and how to mitigate it. Tooling: Practical experience across the current stack , e.g., Braintrust (CI regression), Arize Phoenix (OpenTelemetry-native observability), Promptfoo (OWASP red-team), Galileo, DeepEval/Confident AI, Ragas, Langfuse/LangSmith , and the judgment to know what each metric proves. Trace Fluency: Comfort reading OpenTelemetry-style agent traces (spans for retrieval, tool calls, sub-agent handoffs) and attributing failures to components. Skeptical Communication: Defends uncomfortable numbers to delivery leadership and client stakeholders; influences engineering without authority.

One address, no account. We’ll tell you when matching roles go live.

More at SIMPLIIGENCE INC

Related open roles

View all roles
LLM Agent Evaluation Engineer at SIMPLIIGENCE INC · Padmi