Can handle a three-part support question by answering the policy and timing subrequests correctly, but it leaves the application-step subrequest unanswered.
What was measured
multi-part query handling
Correctly answers a single support conversation that contains multiple related questions or requests.
transformation
What was given, what came back
Test input: Refund policy, application, and processing-time question · text · group: stylenova-customer-support
Input — what we sent
The exact prompt
What's your refund policy? And if I'm eligible, how do I actually apply for one? Also, how long does it usually take to process?
A multi-part customer-support query asking for the refund policy, how to apply for a refund, and how long processing usually takes. It tests grounded FAQ answering and multi-step policy completeness.
Why this input is hard
- · knowledge-base grounding
- · multi-part query handling
- · policy accuracy
- · response completeness
- · hallucination resistance
Output — unretouched

Also checked on this input — same tool, 4 other criteria
knowledge-bound reasoning✓ WorkedBases the answer on the supplied knowledge base instead of inventing refund-application steps that are not provided.policy accuracy✓ WorkedReturns the correct refund-policy facts: a 30-day return window, unused/unwashed/original-packaging requirements, and a 5–7 business-day refund processing time.response completeness⚠ StruggledLeaves one of the user’s requested items unanswered by not explaining how to apply for a refund, so the support reply is incomplete.uncertainty acknowledgment✓ WorkedClearly acknowledges a knowledge gap by saying it does not have instructions in the knowledge base for how to apply for a refund specifically.
Provenance
- Observation
- 9e3c8999-f97c-4a45-81ae-fe92232da074
- Evidence run
- cf8842ae-b5a5-4bf3-a6d9-d61ebe406097
- Study
- Automate customer support using an AI chatbot
- Research task
- 86b9jm3ev
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "fs-agent",
scenario: "stylenova-customer-support"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on multi-part query handling
CoSupport.ai✓ WorkedThe assistant can resolve a single support message with two linked requests in one reply, covering both the return policy and the delivery timeline; in the test it gave the 30-day return window plus standard shipping at 5-7 business days and express shipping at 2-3 business days for Plus/Elite members.Freshdesk✓ WorkedHandles a single refund question that bundles three related requests by answering them in one numbered, structured reply instead of splitting or dropping sub-questions.JotForm✓ WorkedHandled a single support turn with three linked questions in one reply, covering the policy, how to apply, and processing time without splitting the conversation apart.Kommunicate✓ WorkedHandles a three-part support question in one coherent reply, covering the refund policy, how to start a return, and the usual processing time in a single response.Tidio✓ WorkedHandles a three-part support question in one reply by answering the policy, the application path, and the expected processing-time question in a single message.Wonderchat✓ WorkedIt answered a single 3-part support question in one response, covering the refund policy, how to apply, and the usual processing time.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com