It correctly recognizes a bot-avoidance request as escalation intent and provides the full human-support channel list, including email, phone, and ticket portal.
What was measured
Response Completeness
Addresses all parts of the user’s support request instead of leaving key questions unanswered.
decisive for this rankingtransformation
A support bot that leaves key questions unanswered is not solving the customer’s issue. (3 of 3 judges)
What was given, what came back
Test input: Customer-support handoff request · text
Input — what we sent
The exact prompt
I received the wrong item in my order. I've already checked the order details and this is clearly a mistake on your end. I don't want any more back and forth — can you please connect me to a customer support agent or raise a ticket for this?
A frustrated support-escalation request after receiving the wrong item, asking to be connected to a support agent or have a ticket raised.
Why this input is hard
- · Human handoff reliability
- · Ticket creation workflow
- · Escalation contextual awareness
- · Professional tone under frustration
Output — unretouched

Also checked on this input — same tool, 3 other criteria
Multilingual Understanding✓ WorkedIt understands the Hindi support request and answers fully in Hindi, preserving the support-channel options and response-time details.Multilingual Understanding✓ WorkedIt understands the Spanish support request and replies fully in Spanish with the same multi-channel escalation options and timing information.Policy Accuracy✓ WorkedIt follows the billing workflow for a suspected double charge by asking whether the entries are pending or posted, requesting transaction references, and reserving escalation for a genuine duplicate posted charge.
Provenance
- Observation
- 9e93e0bf-2bf4-431c-82d4-0027fab38b30
- Evidence run
- 2645dc92-49df-478a-b809-21dfd09f06a7
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "wonderchat"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 5 other tools
measured on Response Completeness
Chatbase✓ WorkedFor a user who does not want to talk to a bot, it asks only for the email address and a brief issue description to create the support request.Freshdesk◐ MixedOn a double-charge fraud claim, the bot does not escalate immediately; it first frames the issue as a possible authorization hold and asks for transaction details, so the handoff is delayed rather than direct.FS Agent✓ WorkedCompletes a 'no bot' escalation by routing the user to support and requesting the email address and brief issue description needed to proceed.Respond✓ WorkedWhen the user said they did not want to talk to a bot, the bot handed the conversation to support and confirmed the connection.Zendesk✓ WorkedThe bot escalates a fraud-like duplicate-charge complaint immediately and does not attempt to troubleshoot the payment itself.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com