Tool in benchmark · Version 1

Basedash in AI Database Agents

Scenario-level performance from current published Results.

No published Results are available for Basedash in this benchmark. · 28 scenarios in the benchmark

Publication availability does not describe testing status.

How Basedash performed

Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.

Analytics & Observability4 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
A user reports a bad answer and the operator has to find itNo published result
Reported analytics match what actually happenedNo published result
Someone needs to see what people have been askingNo published result
The operator needs to find the questions the agent is failing onNo published result
Answer Presentation4 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
The answer is a plain fact or a short listNo published result
The answer is a trend or a comparisonNo published result
The answer needs a summary and its detail togetherNo published result
The user asks to see the same answer a different wayNo published result
Conversational Interaction6 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
Customer corrects information given earlierNo published result
The agent offers what to ask nextNo published result
The follow-up is ambiguousNo published result
The next question refers to the previous answerNo published result
The user changes direction mid-conversationNo published result
The user narrows what they just askedNo published result
Data Access Control3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
An excluded column is asked for directlyNo published result
An excluded table is needed to answer the questionNo published result
Two users are given different data scopeNo published result
Question Answering4 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
The answer is available in the connected dataNo published result
The connected data cannot answer the questionNo published result
The question uses a term the company defines itselfNo published result
The user's wording doesn't match how the values are storedNo published result
Reporting3 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
The report is run again after the data has changedNo published result
The user wants several findings in one placeNo published result
The user wants to keep an answer they just gotNo published result
Training4 scenarios · 0 with published Results
ScenarioPublished outcomesTest coverageResult
The agent used the wrong definition and the user corrects itNo published result
The database's names are cryptic and the user documents themNo published result
The user sets a standing rule for all future answersNo published result
The user supplies example questions and the queries they trustNo published result

Reading these Results

Published evidence and test coverage answer different questions.

Publication availability

Which scenarios have a Result?

A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.

Test coverage

What does each Result cover?

Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.

Scenario scope

Inventory is not testing progress

The 28 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Basedash.

Basedash in AI Database Agents | AI Demos