The response was concise and support-appropriate, but the report notes it stayed as one unformatted paragraph with no headers or bullets, which made it harder to scan.
What was measured
conversational quality
Responds in a clear, natural, and support-appropriate way.
transformation
What was given, what came back
Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?
A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.
Why this input is hard
- · knowledge-base grounding
- · multi-part query handling
- · policy accuracy
- · response completeness
- · hallucination resistance
Output — unretouched

Also checked on this input — same tool, 3 other criteria
multi-part query handling✓ WorkedHandled a single support turn with three linked questions in one reply, covering the policy, how to apply, and processing time without splitting the conversation apart.policy accuracy✓ WorkedMatched the refund rules exactly: a 30-day return window, unused/unwashed items in original packaging, Elite members with free returns, $4.99 return shipping for others, and refunds processed in 5-7 business days.response completeness✓ WorkedAnswered all requested parts of the refund request, including eligibility, application steps, and the 5-7 business day processing estimate, and also offered to guide the return process.
Provenance
- Observation
- fb9d3df9-cc07-4e78-9403-fd92b7fb0b8c
- Evidence run
- cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "jotform",
scenario: "stylenova-customer-support"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on conversational quality
Freshdesk◐ MixedKeeps the answer clearly structured and support-appropriate, but the researcher notes it is somewhat long and text-heavy.Kommunicate✓ WorkedKeeps the refund answer concise and easy to read, presenting the information in a single support-friendly paragraph.Tidio✓ WorkedResponds in a clear, support-appropriate format with bold section headers and bullet points that are easy to scan.Wonderchat◐ MixedThe response was clearly structured and support-appropriate, but the report says it was slightly over-explained for a simple policy check and could overwhelm a customer who wanted a quick answer.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com