Demo built for Maven AGI by The AI Pipe theaipipe.com
Four axes, graded separately. One stale article.
Forty annotated support conversations run against a live agent with real tool use. Answer quality,
retrieval quality, action correctness and escalation behavior are scored by four independent graders,
and the run is compared against the last recorded nightly to separate a regression from a known failure.
This morning a knowledge base publish restored one article to a revision written under the old refund
policy. Every reply it produces is fluent, faithful to its source, and authorises money the company
does not owe. Only one of the four axes notices.
“Design and run tests and evals for AI behavior, including answer quality, retrieval quality,
action correctness, escalation behavior, and regression risks.”
Senior Forward Deployed Engineer, Maven AGI. This bench is that sentence, built.
Agent under test
Name…
Model…
Tools…
Suite
Conversations…
Knowledge base…
Orders and accounts…
Baseline
Run…
Recorded…
Known failures…
Knowledge base publish
Loading.
Loading the suite.
CaseConversationCategoryAQ / RQ / AC / EBResult
Run the campaign to fill this in.
Send the agent something nobody scripted
Write a customer message of your own. It runs against the same agent, the same knowledge base and the
same tools, live, as Cordelia Fenwick (CUS-2001), a Priority Care member whose queen was delivered 240
nights ago on order AS-79114. The bench has no label for your message, so it grades the reply on answer
quality and shows you what the agent did to the account. Compare the two.