Someone needs to see what people have been asking
A usage view shows what people asked and attributes those questions correctly.
What this scenario means
This scenario checks whether the product can surface a usage view that reflects real demand, not just raw logs. A good agent shows the questions people asked, keeps them attributable, and accounts for the full set without dropping, merging, or mislabelling entries. That matters because oversight depends on seeing actual usage clearly.
What we evaluate
- Whether the usage view lists the questions that were asked.
- Whether each question is attributed to the source that asked it.
- Whether the view accounts for the full set of questions, not just a subset.
- Whether the totals or counts shown by the view match the underlying activity.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Analytics & Observability
Lets whoever runs the agent see what people actually asked and what actually happened — and reports numbers that match reality. ON TRIAL: aggregate analytics are S57 and S24; individual inspection is S58 and S59. Demotes if the scenarios cannot discriminate. Deliberately NOT the same capability as the Customer Support benchmark's 'Analytics and observability' (C8) — there the operator is a different person from the user; here the buyer IS the user and sees every answer (rule C-10).
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
AI Database Agents
Which AI database agent answers business questions about a live relational database correctly — in the company's own terms, and honestly when the data cannot answer?