Managed RAG
Managed RAG evaluates services that turn private documents into a query API, so buyers can compare buying a managed service with building the pipeline themselves across grounding, citation and refresh behaviour.
What this benchmark is
This benchmark covers managed RAG and RAG-as-a-Service products that let a buyer upload or connect private documents and get a query API back. It includes services that only retrieve passages and services that also generate answers.
The buying question is build versus buy: whether a managed service is better than assembling and running the retrieval pipeline yourself. Readers learn whether the service makes private knowledge reachable, stays grounded in what it retrieved, cites the supporting source, and catches up when the source changes.
The benchmark uses fictional documents and the same documents and questions for every tool. One fixed document set is used for ingestion, retrieval, grounded generation and citations; a separate document set is changed for Knowledge Sync so freshness testing does not contaminate the others. Each tool runs on its own default setup, and a build-it-yourself pipeline is included as a subject for the build-versus-buy comparison.
In scope
- Services where you upload or connect documents and get a query API back.
- Services that only retrieve and services that also generate answers.
- Connectors and automatic updates.
- Retrieval machinery sold as a service, including hybrid search, reranking and metadata filters.
Out of scope
- Raw vector databases.
- Website chat widgets.
- Web search.
- Document parsing on its own.
- Multi-tenant isolation and per-user permissions.
- Structured Document Extraction's querying over already extracted records; here Retrieval works on document text.
Capabilities included
The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.
| Capability | What it means here | Scenarios |
|---|---|---|
| Data Ingestion | The handed-over files become reachable, and the service reports when ingestion did not work. | 4 |
| Retrieval | The passage that answers the question comes back from the documents when it exists, and a missing answer is discoverable when it does not. | 4 |
| Grounded Generation | The answer stays faithful to the passages the system itself retrieved, including declining when those passages do not support an answer. | 3 |
| Citations | The cited passage genuinely supports the claim the answer makes, even when the answer itself is wrong. | 1 |
| Knowledge Sync | The indexed state catches up with the source after content is added, changed, or removed. | 3 |
Scenarios included
A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.
Data Ingestion
4 scenariosRetrieval
4 scenariosGrounded Generation
3 scenariosHow the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Unchanged vendor default
Every tool is tested on its own default setup, unchanged.
Build-it-yourself control
A build-it-yourself pipeline runs as a subject alongside the managed services.
Judge retrieval against the documents
Retrieval is judged against the documents: the supporting passage must come back when it exists.
Judge generation against retrieved material
Grounded generation is judged against what the tool actually retrieved, not against the ideal answer.
Do not score one miss twice
A retrieval miss is not charged again in generation; declining is correct when the retrieved material does not support an answer.
Judge citations by support
Citations are judged by whether the cited passage supports the claim, not by whether the answer is correct.
Correct citation, wrong answer
A correct citation can point to the source of a wrong answer; a citation attached to an invented claim is aggravating.
Unsupported capabilities score zero
A capability a tool does not support scores 0, and the 0 stays in the denominator.
End-to-end answer correctness is separate
End-to-end answer correctness is reported separately as a composite outcome.
Follow the documented refresh path
Each Knowledge Sync test follows the tool's own documented update path to completion; time is recorded, and non-completion counts as non-convergence.
Scoring is still open
Scoring, weighting, critical failures, the eligibility contract, and corpus scale are still open methodology work.
V1 is accepted, but scoring and scenario rubrics are still open.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
Grounded QA fixture corpora — Fenwake documents for the Managed RAG ranking
The registered fixture corpus supplies the fictional documents used by the Managed RAG ranking.