Source description
About the role
We are looking for an Evaluation Engineer to own and scale the evaluation program that gates our LLM and agentic system releases. Please share an updated profile to kavitha@simpliigence.com (+91) 74839 25904 Mode of Work: Remote Required Skills & Experience Eval Ownership: 6+ years software/ML engineering with demonstrated ownership of an LLM/agent evaluation program that gated real releases. Strong Python (pandas, SQL, pytest). Agent (not just LLM) Evals: Hands-on with trajectory/trace-based evaluation, tool-calling and planning metrics, and pass^k multi-run methodology , output-only eval experience is not enough. Judge Validation: Proven ability to build and statistically validate LLM-as-judge / agent-as-judge evaluators and to articulate why naive judges fail on multi-turn traces and how to mitigate it. Tooling: Practical experience across the current stack , e.g., Braintrust (CI regression), Arize Phoenix (OpenTelemetry-native observability), Promptfoo (OWASP red-team), Galileo, DeepEval/Confident AI, Ragas, Langfuse/LangSmith , and the judgment to know what each metric proves. Trace Fluency: Comfort reading OpenTelemetry-style agent traces (spans for retrieval, tool calls, sub-agent handoffs) and attributing failures to components. Skeptical Communication: Defends uncomfortable numbers to delivery leadership and client stakeholders; influences engineering without authority. We are looking for an Evaluation Engineer to own and scale the evaluation program that gates our LLM and agentic system releases. Please share an updated profile to kavitha@simpliigence.com (+91) 74839 25904 Mode of Work: Remote Required Skills & Experience Eval Ownership: 6+ years software/ML engineering with demonstrated ownership of an LLM/agent evaluation program that gated real releases. Strong Python (pandas, SQL, pytest). Agent (not just LLM) Evals: Hands-on with trajectory/trace-based evaluation, tool-calling and planning metrics, and pass^k multi-run methodology , output-only eval experience is not enough. Judge Validation: Proven ability to build and statistically validate LLM-as-judge / agent-as-judge evaluators and to articulate why naive judges fail on multi-turn traces and how to mitigate it. Tooling: Practical experience across the current stack , e.g., Braintrust (CI regression), Arize Phoenix (OpenTelemetry-native observability), Promptfoo (OWASP red-team), Galileo, DeepEval/Confident AI, Ragas, Langfuse/LangSmith , and the judgment to know what each metric proves. Trace Fluency: Comfort reading OpenTelemetry-style agent traces (spans for retrieval, tool calls, sub-agent handoffs) and attributing failures to components. Skeptical Communication: Defends uncomfortable numbers to delivery leadership and client stakeholders; influences engineering without authority.
More at SIMPLIIGENCE INC
Related open roles
Salesforce Qualtrics Integration Specialist
Bangalore
Senior Business Analyst(Healthcare - Payer Domain) (Bengaluru)
Bangalore
Senior Business Analyst(Healthcare - Payer Domain)
India
Technical Lead Salesforce Commerce Cloud
India
Sr Senior Salesforce Developer
Bangalore
Sr Senior Salesforce Developer
Bangalore