Capability in benchmark · Version 1

Training

Uses company-specific information to improve later behaviour and to generalise beyond the exact example given.

4 scenarios · 2 current published Results · 2 tools with published Results

How the tools performed

Explore current published Results across this capability’s scenarios. Open a scenario to compare its tools, or a Result to inspect the evidence.

Each cell reports its own published test set. There is no combined capability score or claim that different Results used identical tests.

Scenarios in this capability
Tools with published Results appear first, alphabetically. Remaining participants stay visible below.
Participating toolThe agent used the wrong definition and the user corrects it →The database's names are cryptic and the user documents them →The user supplies example questions and the queries they trust →The user sets a standing rule for all future answers →
BlazeSQLThe agent used the wrong definition and the user corrects it →No published resultThe database's names are cryptic and the user documents them →
1 Pass
1/1 assessed1/1 gradableView Result →
The user supplies example questions and the queries they trust →No published resultThe user sets a standing rule for all future answers →No published result
FutureSmart Database AgentThe agent used the wrong definition and the user corrects it →No published resultThe database's names are cryptic and the user documents them →No published resultThe user supplies example questions and the queries they trust →No published resultThe user sets a standing rule for all future answers →
1 Pass
1/1 assessed1/1 gradableView Result →
AI for DatabaseNo published result across these scenarios
Anomaly AINo published result across these scenarios
AskYourDatabaseNo published result across these scenarios
BasedashNo published result across these scenarios
camelAINo published result across these scenarios
DefiniteNo published result across these scenarios
DotNo published result across these scenarios
DraxlrNo published result across these scenarios
QuerioNo published result across these scenarios

Reading the comparison

Publication availability and test coverage describe different things.

Published evidence

Publication is not testing progress

“No published result” says only that there is no current public Result. It does not indicate whether a tool has been tested, passed, failed or is applicable.

Test coverage

Read each Result’s test set

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Each fraction uses the pinned tests in that same published Result.

Go deeper

Choose the view you need

A scenario compares tools on one situation. A tool page follows one tool across this benchmark. A Result explains the outcome and shows the evidence.

Training in AI Database Agents | AI Demos