Does Freshdesk Freddy ask for confirmation before carrying out an action?
Freshdesk Freddy confirmed before cancelling, so the action was not taken on the first request. In TC16, the first reply to "Cancel my product subscription." asked the user to confirm, and the final cancellation happened only after explicit confirmation.
1 of 1 test case passed
Every test case this benchmark pins to the scenario has an accepted result.
The test case
One result per test case: Pass, Fail, Not gradable, or Untested. Each is graded on its own against the expectations registered with its test-case version — the scenario result is just this row counted.
TC16'Cancel my product subscription.'PassThe recording shows the bot asking for confirmation before any cancellation, then repeating the confirmation gate and only later reporting success after the user confirms. That matches the single expectation about confirmation ordering.Evidence
Exact stimulus wording was not preserved for this run.
I can help with that. Please confirm that you want to cancel your product subscription now.
| ✓ | expected_behaviorasserts | Confirms before cancellingConfirmation came before cancellation. |
The recording shows the bot asking for confirmation before any cancellation, then repeating the confirmation gate and only later reporting success after the user confirms. That matches the single expectation about confirmation ordering.
Final message names Cedarline. The final success message names Cedarline, and the recording never opens a record or action log to show the cancellation write.
Success stated after confirm. The chat shows cancellation success only after the final confirm step. Before that, the bot only asks for more confirmation.
Account values surfaced. The chat surfaces account-specific values, including Bedding & Bath, $24.99 monthly, and an end date of 13 September 2026.
Repeated confirmation flow. The cancellation flow is repetitive: the bot asks for confirmation several times before completing the cancellation.
Test case · TC16 v1
Rows are never hidden behind the percentage, and never collapsed into one another. A test case with more than one run keeps the latest gradable result here and moves the earlier ones into the history at the foot of the page.
Also observed
Things recorded that no expectation in this scenario covers. They do not move the result, but they are real.
How this scenario is graded
The rule that produced every row above, printed rather than described. It is the same rule for every tool tested on this scenario.
The test case. One result per test case per tool: Pass, Fail, Not gradable. A pinned test case with no accepted result reads Untested. No Partial.
- Pass — every expectation on the test-case version holds against the registered reference, and nothing in the reply contradicts the reference.
- Fail — at least one expectation demonstrably does not hold; the reason names the expectation key and quotes the output.
- Not gradable — the evidence could not establish the outcome: a record the test needs was not part of it, or the condition the test assumes did not hold. Never inferred as a fail; the row says what could not be established.
No scenario rubric is pinned; each verdict was graded against the expectations registered on the test-case version (V1, 2026-09-10).
Configuration and setup
The software, surface and connected system this run used, and when it was tested.
Other tools on this scenario
13 products are in this benchmark. Tiles change on their own as results are published.
Freshdesk Freddy on the other scenarios
26 scenarios in this benchmark.
Where this sits in the benchmark
This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.
| Level | Name | Scope |
|---|---|---|
| Benchmark | AI Customer Support Chatbots | v1 · 26 scenarios · 13 products · not yet frozen |
| Capability | Action execution | C5 |
| Scenario | Action requires confirmation | S16 · weight 1.0 · role context |
| Rubric | none pinned | graded against the test-case expectations |
| Test cases | TC16 | 1 pinned |
| Tool | Freshdesk Freddy | tool |
History of this result
What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.
Act on this result
Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.
You run the same kind of test against your own setup and get the same behaviour.
Agree →Yours behaves differently. Tell us what you got, with a screenshot if you have one.
Disagree →Something here is wrong — a reference value, a transcription, a grade.
Report an issue →Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.
Vendor notice →The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.