On ambiguity queries, the Search API reached 71% top-3, indicating it usually resolves entity collisions instead of drifting to the wrong interpretation.
What was measured
Ambiguity handling
Does it resolve entity collisions or return the wrong "Apollo"?
decisive for this rankingtransformation
Disambiguating entities and returning the right result is fundamental to search quality, especially for agent workflows. (3 of 3 judges)
What was given, what came back
Input — what we sent
No input — this is a capability finding
The observation is about the tool itself rather than one test input, so there is nothing to show on this side by design.
Provenance
- Observation
- b72846f9-2723-4e84-8056-e377cb549fdd
- Evidence run
- 5f092d61-1d8a-457f-8b1d-19b45e17ef0d
- Study
- Search the Web From an Agent — Web Search & Answer APIs Compared
- Research task
- 86bazaep1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- output only
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "you-com"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Ambiguity handling
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com