Covers all requested parts of the refund flow, including eligibility, how to create a return, and the timing/process guidance, rather than stopping after the policy summary.
What was measured
response completeness
Addresses all parts of the user’s support request instead of leaving key questions unanswered.
transformation
What was given, what came back
Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?
A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.
Why this input is hard
- · knowledge-base grounding
- · multi-part query handling
- · policy accuracy
- · response completeness
- · hallucination resistance
Output — unretouched

Also checked on this input — same tool, 3 other criteria
conversational quality◐ MixedKeeps the answer clearly structured and support-appropriate, but the researcher notes it is somewhat long and text-heavy.multi-part query handling✓ WorkedHandles a single refund question that bundles three related requests by answering them in one numbered, structured reply instead of splitting or dropping sub-questions.policy accuracy✓ WorkedStays aligned with the stated refund policy: returns within 30 days, items unused/unwashed/in original packaging, exchanges for size or color, free returns for Elite members, and $4.99 shipping for other customers.
Provenance
- Observation
- 751b82d7-de72-4846-9b38-14a702e232b4
- Evidence run
- cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "freshdesk",
scenario: "stylenova-customer-support"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on response completeness
CoSupport.ai✓ WorkedIt answers every part of the customer’s question instead of dropping a sub-request: the reply includes the return rule, the standard delivery estimate, and the express-delivery estimate.FS Agent⚠ StruggledLeaves one of the user’s requested items unanswered by not explaining how to apply for a refund, so the support reply is incomplete.JotForm✓ WorkedAnswered all requested parts of the refund request, including eligibility, application steps, and the 5-7 business day processing estimate, and also offered to guide the return process.Kommunicate◐ MixedPartially answers a multi-part refund request: it gives the policy and timing, but the application procedure stays vague and the reply omits Elite free returns and the $4.99 non-Elite return-shipping rule that the report says were in the knowledge base.Tidio✓ WorkedFully covers all requested parts of the support request instead of leaving a sub-question unanswered.Wonderchat✓ WorkedIt fully covered the refund request by giving the policy, the application steps, the processing timeline, and contact options for starting the return.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com