Stays aligned with the stated refund policy: returns within 30 days, items unused/unwashed/in original packaging, exchanges for size or color, free returns for Elite members, and $4.99 shipping for other customers.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedFreshdesk
What was measured
policy accuracy

Gives policy details that match the provided knowledge base and does not misstate rules.

transformation

What was given, what came back

Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?

A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.

Why this input is hard
  • · knowledge-base grounding
  • · multi-part query handling
  • · policy accuracy
  • · response completeness
  • · hallucination resistance
Output — unretouched
image
Provenance
Observation
2519cd47-ddf2-4f14-8651-3767ebf80e64
Evidence run
cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
Study
Automate customer support using an AI chatbot
Research task
86b9jm3ev
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "freshdesk",
  scenario: "stylenova-customer-support"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on policy accuracy
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com