The user wants several findings in one place
A user asks to combine several distinct findings into one report-like artifact.
What this scenario means
This scenario tests whether the agent can assemble more than one answer into a single coherent output without breaking any part of the underlying results. A good agent keeps each finding tied to the question that produced it, preserves consistency across the combined artifact, and does not introduce new analysis just because the findings are shown together.
What we evaluate
- Whether the agent combines several distinct findings into one coherent artifact.
- Whether each included finding still matches the answer to its own original question.
- Whether the combined output keeps the findings consistent with one another rather than mixing, dropping, or altering them.
- Whether the result remains a report-like summary of already produced answers, not a fresh analysis with changed meaning.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Reporting
Turning analysis the agent produced into a durable, re-runnable artifact that still means the same thing later.
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
AI Database Agents
Which AI database agent answers business questions about a live relational database correctly — in the company's own terms, and honestly when the data cannot answer?