Benchmark · Version 1

Web Search for AI

Evaluates whether live-web search and answer APIs find supporting material, notice when it is absent, keep up with web changes, and stay grounded in their own sources when they answer.

What this benchmark is

This benchmark looks at live-web search and answer products used by AI applications. It asks whether a product can find web material that supports an answer, keep pace with change, and stay grounded in its own sources when it answers.

A reader learns how these products behave when support exists, when it is missing, when pages change, and when sources disagree. It is about the buyer question, not about a particular corpus or interface.

The benchmark tests the single-call search layer of live-web search and answer APIs, including SERP-data APIs as subject types. It works on the open web rather than uploaded corpora, and it keeps Web Page → Markdown and multi-step research endpoints outside the boundary.

In scope

  • Live-web search and answer APIs used by AI applications.
  • Products that must find the URL rather than be given it.
  • SERP-data APIs as first-class subjects.
  • The single-call search layer.
  • Live-web source material that includes titles, snippets and answer boxes.

Out of scope

  • Web Page → Markdown.
  • Managed RAG.
  • Multi-step research endpoints.
  • Structural fidelity grading.
3Capabilities
10Scenarios

Capabilities included

The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.

CapabilityWhat it means hereScenarios
Web RetrievalWeb Retrieval means the product can find live-web material that supports an answer when it exists, and can tell when it does not.4
FreshnessFreshness means the product's results match the web as it is now, including new pages, changed pages, and removed pages.3
Answer GenerationAnswer Generation means that, where the product writes the answer, it says only what retrieved sources support, declines when support is absent, flags disagreement, and attributes claims.3

Scenarios included

A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Endpoint boundary

Single-call search layer

This benchmark tests the single-call search layer, not a multi-step research endpoint.

Subject class

SERP-data APIs count

SERP-data APIs are first-class subjects, and their titles, snippets and answer boxes are treated as source material.

Separate boundary

Web Page → Markdown stays separate

Web Page → Markdown is a different boundary where the URL is already known, and structural fidelity is not graded here.

Answer grounding

Source-limited answers

When a product writes the answer, it should say only what retrieved sources support, decline when support is absent, flag source disagreement, and attribute claims.

Freshness method

Probe before and after

Freshness runs use a before probe and an after probe; if the before probe comes out the wrong way, the run is void, not failed.

Method status

Ranking and scoring remain undesigned

Ranking and scoring are in scope but remain undesigned.

Ranking and scoring are in scope but remain undesigned.

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Not yet written. No registered resource is declared on this benchmark version or required by its pinned test cases.
Web Search for AI — Benchmark definition | AI Demos