Fully covers all requested parts of the support request instead of leaving a sub-question unanswered.
What was measured
response completeness
Addresses all parts of the user’s support request instead of leaving key questions unanswered.
transformation
What was given, what came back
Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?
A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.
Why this input is hard
- · knowledge-base grounding
- · multi-part query handling
- · policy accuracy
- · response completeness
- · hallucination resistance
Output — unretouched

Also checked on this input — same tool, 3 other criteria
conversational quality✓ WorkedResponds in a clear, support-appropriate format with bold section headers and bullet points that are easy to scan.multi-part query handling✓ WorkedHandles a three-part support question in one reply by answering the policy, the application path, and the expected processing-time question in a single message.policy accuracy✓ WorkedReturns policy details that match the described support policy: returns within 30 days of delivery, items must be unused/unwashed/in original packaging, Elite members get free returns, and all others pay a $4.99 return-shipping fee.
Provenance
- Observation
- 75d31afd-23c8-4584-b77f-d51cbf42b44e
- Evidence run
- cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "tidio",
scenario: "stylenova-customer-support"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on response completeness
CoSupport.ai✓ WorkedIt answers every part of the customer’s question instead of dropping a sub-request: the reply includes the return rule, the standard delivery estimate, and the express-delivery estimate.Freshdesk✓ WorkedCovers all requested parts of the refund flow, including eligibility, how to create a return, and the timing/process guidance, rather than stopping after the policy summary.FS Agent⚠ StruggledLeaves one of the user’s requested items unanswered by not explaining how to apply for a refund, so the support reply is incomplete.JotForm✓ WorkedAnswered all requested parts of the refund request, including eligibility, application steps, and the 5-7 business day processing estimate, and also offered to guide the return process.Kommunicate◐ MixedPartially answers a multi-part refund request: it gives the policy and timing, but the application procedure stays vague and the reply omits Elite free returns and the $4.99 non-Elite return-shipping rule that the report says were in the knowledge base.Wonderchat✓ WorkedIt fully covered the refund request by giving the policy, the application steps, the processing timeline, and contact options for starting the return.
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com