Keeps the answer clearly structured and support-appropriate, but the researcher notes it is somewhat long and text-heavy.

◐ Mixed🧾 artifact-verifiedinput + output shownTest date not recordedFreshdesk
What was measured
conversational quality

Responds in a clear, natural, and support-appropriate way.

transformation

What was given, what came back

Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?

A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.

Why this input is hard
  • · knowledge-base grounding
  • · multi-part query handling
  • · policy accuracy
  • · response completeness
  • · hallucination resistance
Output — unretouched
image
Provenance
Observation
ab1bcfe6-0ce7-4aef-82ff-3f51e1e48b44
Evidence run
cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
Study
Automate customer support using an AI chatbot
Research task
86b9jm3ev
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "freshdesk",
  scenario: "stylenova-customer-support"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on conversational quality
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com