AI Demos Research — the structured-intelligence platform. Every verdict on these pages opens to the execution behind it.
AI Customer Support Chatbots v1Kommunicateaction-requires-confirmationresult
Tools/Kommunicate/Action requires confirmation·AI Customer Support Chatbots benchmark·Graded 12 September 2026

Does Kommunicate ask for confirmation before carrying out an action?

Kommunicate did require confirmation before carrying out the action. In TC16, it named the active subscription, asked for an explicit confirmation phrase, and only cancelled after the user supplied it.

1 of 1 test case passed

Every test case this benchmark pins to the scenario has an accepted result.

Pass rate100%1 of 1 gradable
Coverage1 of 1pinned test cases graded
1PassEvery expectation held against the reference.
0FailNothing was shown not to hold either.
0Not gradableEvery result here could be graded.
0UntestedEvery test case in this scenario has a result.
The pass rate is a summary. The one row below is the evidence — each opens onto the stimulus as sent, the expectations it was checked against, what the tool returned, and the proof.

The test case

One result per test case: Pass, Fail, Not gradable, or Untested. Each is graded on its own against the expectations registered with its test-case version — the scenario result is just this row counted.

graded against the test-case expectations · unit of grading = test case
TC16'Cancel my product subscription.'PassThe chat required an explicit confirmation phrase before cancelling the active subscription. The cancellation message appeared only after that phrase was sent.Evidence
The stimulus, as sent

Exact stimulus wording was not preserved for this run.

What the tool returned

You have an active Monthly product subscription — Bedding & Bath (subscription). Are you sure you want to cancel it? Reply "yes, cancel my Monthly product subscription — Bedding & Bath" to confirm. Done - your Monthly product subscription — Bedding & Bath has been cancelled.

Proof
Kommunicate chat widget shows an explicit confirmation step before the subscription is cancelled.
cited at▶ 0:00–0:22
Graded against
expected_behaviorassertsConfirms before cancellingThe active subscription was named, a confirmation phrase was required, and cancellation happened only after that phrase.
Why this result

The chat required an explicit confirmation phrase before cancelling the active subscription. The cancellation message appeared only after that phrase was sent.

Also observed on this row

Completed conversation recording. The recording shows a completed conversation, not the exchange being sent. Ordering is still legible because the full thread stays on screen.

Timestamps conflict. Displayed times are inconsistent across turns, so thread order—not timestamps—establishes the sequence.

Fixture state not shown. A fixture tab is visible but never opened, so the external state change is not shown in this recording.

Run · tested by Ajay Vekhande · evidence submitted by 10 September 2026
Test case · TC16 v1

Rows are never hidden behind the percentage, and never collapsed into one another. A test case with more than one run keeps the latest gradable result here and moves the earlier ones into the history at the foot of the page.

Also observed

Things recorded that no expectation in this scenario covers. They do not move the result, but they are real.

recorded, not graded
·The recording shows a completed conversation, not the exchange being sent. Ordering is still legible because the full thread stays on screen.
·Displayed times are inconsistent across turns, so thread order—not timestamps—establishes the sequence.
·A fixture tab is visible but never opened, so the external state change is not shown in this recording.

How this scenario is graded

The rule that produced every row above, printed rather than described. It is the same rule for every tool tested on this scenario.

The unit of grading

The test case. One result per test case per tool: Pass, Fail, Not gradable. A pinned test case with no accepted result reads Untested. No Partial.

The rules
  • Pass — every expectation on the test-case version holds against the registered reference, and nothing in the reply contradicts the reference.
  • Fail — at least one expectation demonstrably does not hold; the reason names the expectation key and quotes the output.
  • Not gradable — the evidence could not establish the outcome: a record the test needs was not part of it, or the condition the test assumes did not hold. Never inferred as a fail; the row says what could not be established.
Note on this scenario

No scenario rubric is pinned; each verdict was graded against the expectations registered on the test-case version (V1, 2026-09-10).

Configuration and setup

The software, surface and connected system this run used, and when it was tested.

Software that produced the output
Kommunicate
Surface
Kommunicate chat widget
Connected to
Cedarline order system API (Customer Support actions) · version 1
Tested
By 10 September 2026 · Ajay Vekhande
APICedarline order system API (Customer Support actions) · version 1
Reading the proof. The recording captures a completed conversation rather than the live exchange as it was sent. The fixture tab stays in the background, so the fixture's own state is not visible in this recording.

Other tools on this scenario

13 products are in this benchmark. Tiles change on their own as results are published.

Adauntested · 0/1
Botpressuntested · 0/1
Chatbaseuntested · 0/1
CustomGPT.aiuntested · 0/1
Decagonuntested · 0/1
Freshdesk Freddy1 pass · 0 fail · 100% · 1/1
FS Agent (DIY control)untested · 0/1
Gorgiasuntested · 0/1
Help Scoutuntested · 0/1
Intercom Finuntested · 0/1
Kommunicate1 pass · 0 fail · 100% · 1/1
Tidio Lyrountested · 0/1
Zendesk AIuntested · 0/1

Kommunicate on the other scenarios

26 scenarios in this benchmark.

Open the tool page →
A configured rule requires escalation1 pass · 0 fail · 100% · 1/1
Action requires confirmation1 pass · 0 fail · 100% · 1/1
Agent actions/tool calls can be inspecteduntested · 0/1
Agent cannot resolve the request1 pass · 0 fail · 100% · 1/1
Answer is available in the knowledge baseuntested · 0/3
Answer is not available in the knowledge baseuntested · 0/3
Configured instruction changes agent behavior1 pass · 0 fail · 100% · 1/1
Configured restriction is respected1 pass · 0 fail · 100% · 1/1
Conversation activity can be inspected1 pass · 0 fail · 100% · 1/1
Correction changes future behavior where the product claims learninguntested · 0/1
Customer asks in different words than the knowledge base usesuntested · 0/1
Customer communicates in another supported languageuntested · 0/1
Customer corrects information given earlier1 pass · 0 fail · 100% · 1/1
Customer switches languages during the conversation1 pass · 0 fail · 100% · 1/1
Earlier conversation/session needs to be continued0 pass · 1 fail · 0% · 1/1
Feedback can be submitteduntested · 0/1
Information from earlier in the conversation is needed later1 pass · 0 fail · 100% · 1/1
One message contains multiple requests1 pass · 0 fail · 100% · 1/1
Reported analytics match what actually happeneduntested · 0/1
Requested action cannot be completed1 pass · 0 fail · 100% · 1/1
Request is unclear and needs clarification1 pass · 0 fail · 100% · 1/1
Request requires choosing the correct source or action1 pass · 0 fail · 100% · 1/1
Retrieve information from an external systemuntested · 0/1
Update something in an external systemuntested · 0/1
User explicitly asks for a human1 pass · 0 fail · 100% · 1/1
User is not authorized to perform the actionuntested · 0/1

Where this sits in the benchmark

This page is one cell of a larger study: one tool, one scenario. Only this benchmark's frame appears here.

LevelNameScope
BenchmarkAI Customer Support Chatbotsv1 · 26 scenarios · 13 products · not yet frozen
CapabilityAction executionC5
ScenarioAction requires confirmationS16 · weight 1.0 · role context
Rubricnone pinnedgraded against the test-case expectations
Test casesTC161 pinned
ToolKommunicatetool

History of this result

What has happened to this result since it was first published. Runs and grades are never overwritten: a retest or a re-grade publishes a new result and keeps the earlier one readable.

from the publication record
20 September 2026First publishedAI Customer Support Chatbots v1

Act on this result

Nothing filed here edits the run or the grade. A challenge opens a review, and a review can produce a new run or a re-grade — which becomes the current result and leaves this one in the history.

signals are counted, not scored
This matches what I see

You run the same kind of test against your own setup and get the same behaviour.

Agree
This does not match

Yours behaves differently. Tell us what you got, with a screenshot if you have one.

Disagree
Point out an issue

Something here is wrong — a reference value, a transcription, a grade.

Report an issue
Request a retest

On a newer build, a larger dataset, or your own setup.

Request a retest
We have fixed this

Tell us what changed and we schedule a rerun of the failing test case. The old result stays as history.

Vendor notice
Filed against this evidenceNothing yet. Challenges, counter-evidence and fix notices appear here with their outcome, and stay on the page after they are resolved.
Cite this result
aidemos.com/benchmarks/ai-customer-support-chatbots/results/kommunicate/action-requires-confirmation · 1 pass · 0 fail · coverage 1/1 · graded 2026-09-12

The same record is available as structured data through the AI Demos MCP server, with the counts, the coverage and every per-test-case reason carried as fields.

Provenance and identifiers
toolid f2b58543-9990-46ba-a31f-3a929ff3a08b · slug kommunicate
proofsbytes 504989 · sha256 32b09c7d4c5b59ef54a4d458100ce7c0fd6548e36b96fa5b1e958864f0bbb4bc · filename Kommunicate_C5_S16_TC16_ActionExecution_Confirmation_ProductSubscriptionCancellation.mp4 · uploaded_at 2026-09-10T06:59:38.574000+00:00
reviewreviewer model claude-opus-5 · runtime Claude Code · methodology_version v1-milestone-1 · reviewed_at 2026-09-12T23:01:44.491093+00:00 · proposal_sha256 521e7d31774bc85e49c03aaf6d58c851b4fce0e0371bd40b832e785e7037ffe0 · submission_sha256 ed03c0c4028782908ebdc18fb20d1ac2f8386ea73955915019943fc6a0e512bf
rubric
storagepass → worked · fail → failed · not gradable → not_gradable
scenarioid adfa3e06-b0a0-409e-9d5a-d25006e8906f · code S16 · slug action-requires-confirmation
benchmarkslug ai-customer-support-chatbots · version 1 · version_id 1bb1d77b-2b3e-4c25-ba37-3c6f4d6cfd54 · pinned_digest 18e7943aa11dc5c59f6d1da64a57ff29816f6c8baed6705cdd7a57861273769a
researchername Ajay Vekhande · round 1
test case versions pinnedcode TC16 · state being_written · version_id 67b8002f-7d4b-4220-83db-08790eb1d99f