Clearly acknowledges a knowledge gap by saying it does not have instructions in the knowledge base for how to apply for a refund specifically.

✓ Worked🧾 artifact-verifiedinput + output shownTest date not recordedFS Agent
What was measured
uncertainty acknowledgment

Clearly indicates when the system does not know or cannot verify a policy outcome.

transformation

What was given, what came back

Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?

A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.

Why this input is hard
  • · knowledge-base grounding
  • · multi-part query handling
  • · policy accuracy
  • · response completeness
  • · hallucination resistance
Output — unretouched
image
Provenance
Observation
650b8b73-2d9e-4d7b-8a04-a2bcf72828d4
Evidence run
cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
Study
Automate customer support using an AI chatbot
Research task
86b9jm3ev
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "fs-agent",
  scenario: "stylenova-customer-support"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on uncertainty acknowledgment

No other tool was measured on this criterion for this input.

Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com