Support Agent Eval Bench

Acme Sleep Co. · support suite 2026.09

Four axes, graded separately. One stale article.

Forty annotated support conversations run against a live agent with real tool use. Answer quality, retrieval quality, action correctness and escalation behavior are scored by four independent graders, and the run is compared against the last recorded nightly to separate a regression from a known failure.

This morning a knowledge base publish restored one article to a revision written under the old refund policy. Every reply it produces is fluent, faithful to its source, and authorises money the company does not owe. Only one of the four axes notices.

“Design and run tests and evals for AI behavior, including answer quality, retrieval quality, action correctness, escalation behavior, and regression risks.” Senior Forward Deployed Engineer, Maven AGI. This bench is that sentence, built.

Agent under test

Name
Model
Tools

Suite

Conversations
Knowledge base
Orders and accounts

Baseline

Run
Recorded
Known failures

Knowledge base publish

Loading.

Loading the suite.
CaseConversationCategoryAQ / RQ / AC / EBResult
Run the campaign to fill this in.

Send the agent something nobody scripted

Write a customer message of your own. It runs against the same agent, the same knowledge base and the same tools, live, as Cordelia Fenwick (CUS-2001), a Priority Care member whose queen was delivered 240 nights ago on order AS-79114. The bench has no label for your message, so it grades the reply on answer quality and shows you what the agent did to the account. Compare the two.