Can handle a three-part support question by answering the policy and timing subrequests correctly, but it leaves the application-step subrequest unanswered.

◐ Mixed🧾 artifact-verifiedinput + output shownTest date not recordedFS Agent
What was measured
multi-part query handling

Correctly answers a single support conversation that contains multiple related questions or requests.

transformation

What was given, what came back

Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?

A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.

Why this input is hard
  • · knowledge-base grounding
  • · multi-part query handling
  • · policy accuracy
  • · response completeness
  • · hallucination resistance
Output — unretouched
image
Provenance
Observation
9e3c8999-f97c-4a45-81ae-fe92232da074
Evidence run
cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
Study
Automate customer support using an AI chatbot
Research task
86b9jm3ev
Tested at
not recorded
Source
first-party
Evidence state
verified
Proof shown
input + output shown
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "fs-agent",
  scenario: "stylenova-customer-support"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on multi-part query handling
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com