When a follow-up phrase could refer either to the full order-stage breakdown or to the immediately preceding pending-paid subresult, it did not ask for clarification and instead chose the narrower pending-paid interpretation.
What was measured
Ambiguity Handling
Does the tool clarify unclear business terms instead of guessing silently?
transformation
What was given, what came back
Test input: Order pipeline breakdown with paid-pending edge case and last-month comparison · text · group: ecommerce-nl2sql-benchmark
Input — what we sent
The exact prompt
How many orders do we have at each stage right now? Follow-up 1: What percentage of our orders were successfully delivered vs cancelled? Follow-up 2: Are there any orders that are pending but already paid? Follow-up 3: Compare that to last month — same breakdown, I want to see if things have improved or got worse.
A deeper operational analysis of current order stages, delivery-versus-cancellation rates, pending-but-paid edge cases, and a month-over-month comparison of the same breakdown.
Why this input is hard
- · Order pipeline analysis
- · Percentage calculation
- · Edge-case detection
- · Payment/order status joins
- · Multi-turn context retention
- · Month-over-month comparison
- · Ambiguity handling for 'same breakdown'
Output — unretouched

Also checked on this input — same tool, 2 other criteria
Business Insight✓ WorkedCan turn numeric results into plain-English takeaways; the delivered-vs-cancelled follow-up automatically summarized the split and framed delivered orders as more common than cancelled orders.Chart / Visualization Support✓ WorkedCan auto-generate a correct chart for a simple month-over-month comparison, showing Last Month at 11 pending-paid orders versus Current Month at 20.
Provenance
- Observation
- a37351f9-c20e-4396-aaac-98fda74b1cb5
- Evidence run
- db2bb5d5-0e0e-4cb3-8d76-3555c45c23cd
- Study
- Query Live Databases Using Plain English with AI
- Research task
- 86b9y6c99
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "draxlr",
scenario: "ecommerce-nl2sql-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 1 other tool
measured on Ambiguity Handling
This evidence is published in
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com