Web Search for AI
Evaluates whether live-web search and answer APIs find supporting material, notice when it is absent, keep up with web changes, and stay grounded in their own sources when they answer.
What this benchmark is
This benchmark looks at live-web search and answer products used by AI applications. It asks whether a product can find web material that supports an answer, keep pace with change, and stay grounded in its own sources when it answers.
A reader learns how these products behave when support exists, when it is missing, when pages change, and when sources disagree. It is about the buyer question, not about a particular corpus or interface.
The benchmark tests the single-call search layer of live-web search and answer APIs, including SERP-data APIs as subject types. It works on the open web rather than uploaded corpora, and it keeps Web Page → Markdown and multi-step research endpoints outside the boundary.
In scope
- Live-web search and answer APIs used by AI applications.
- Products that must find the URL rather than be given it.
- SERP-data APIs as first-class subjects.
- The single-call search layer.
- Live-web source material that includes titles, snippets and answer boxes.
Out of scope
- Web Page → Markdown.
- Managed RAG.
- Multi-step research endpoints.
- Structural fidelity grading.
Capabilities included
The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.
| Capability | What it means here | Scenarios |
|---|---|---|
| Web Retrieval | Web Retrieval means the product can find live-web material that supports an answer when it exists, and can tell when it does not. | 4 |
| Freshness | Freshness means the product's results match the web as it is now, including new pages, changed pages, and removed pages. | 3 |
| Answer Generation | Answer Generation means that, where the product writes the answer, it says only what retrieved sources support, declines when support is absent, flags disagreement, and attributes claims. | 3 |
Scenarios included
A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.
Web Retrieval
4 scenariosFreshness
3 scenariosHow the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Single-call search layer
This benchmark tests the single-call search layer, not a multi-step research endpoint.
SERP-data APIs count
SERP-data APIs are first-class subjects, and their titles, snippets and answer boxes are treated as source material.
Web Page → Markdown stays separate
Web Page → Markdown is a different boundary where the URL is already known, and structural fidelity is not graded here.
Source-limited answers
When a product writes the answer, it should say only what retrieved sources support, decline when support is absent, flag source disagreement, and attribute claims.
Probe before and after
Freshness runs use a before probe and an after probe; if the before probe comes out the wrong way, the run is void, not failed.
Ranking and scoring remain undesigned
Ranking and scoring are in scope but remain undesigned.
Ranking and scoring are in scope but remain undesigned.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.