Source description
About the role
Drive high-impact AI initiatives within the Caseware platform transformation — shape priorities and sequencing for your workstream, grounded in user research, domain expertise, and a clear view of technical feasibility.
Design and run rigorous eval frameworks: build golden datasets, define task taxonomies, write rubrics that score AI outputs the way a senior practitioner would, and run structured eval cycles before and after every major release.
Work daily with engineering to translate precise problem statements into well-scoped product requirements — including clear acceptance criteria and eval bars before development begins.
Embed with domain experts and partner accounting and audit firms to validate AI outputs against professional standards — not just user preferences, but what a qualified practitioner would actually sign off on.
Drive continuous discovery with practitioners: shadow workflows, test prototypes in context, and turn 'this doesn't feel right' into specific, actionable failure modes.
Stay ahead of the AI landscape — monitor model developments, new architectures, emerging agent patterns, and competitor moves, and bring a clear point of view on what matters for Caseware.
Define and own product health metrics for AI features: task completion, correction rates, override frequency, trust signals, and time-to-completion — and use them to drive product decisions.
Synthesize signals across engineering, partners, domain experts, customer success, and leadership into a coherent product direction — and communicate it with clarity at every level of the organization.
More at Caseware