The user's wording doesn't match how the values are stored
A user’s term and the stored value mean the same thing, but use different words.
What this scenario means
This tests whether the agent can bridge a synonym gap between the user's wording and the database's stored labels. A good agent recognises the intended category, resolves it to the matching stored value, and answers from the matched records. It should not treat the wording as a mismatch or fall back to an empty result when the data is present.
What we evaluate
- Whether it maps a natural synonym to the stored value it represents.
- Whether it returns the matching records rather than treating the question as unanswered.
- Whether it treats the wording difference as a translation problem, not as missing data.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Question Answering
Answers a question about the connected data correctly — finding the right data, using the company's meaning, matching the user's wording to what is stored — and says so when the data cannot answer.
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
AI Database Agents
Which AI database agent answers business questions about a live relational database correctly — in the company's own terms, and honestly when the data cannot answer?