An ordinary digital text document
A plain digital text document used as the baseline text-only document condition.
What this scenario means
This scenario checks the simplest document condition: ordinary digital prose with no special layout or embedded media. It is hard because there is little to mask small failures, so any omission, garbling, reordering, or invention is easy to miss unless the text is preserved exactly. A good system keeps the prose complete, unchanged, and clearly recognisable as plain text.
What we evaluate
- Whether every paragraph is preserved complete and unchanged.
- Whether nothing is silently missing, garbled, or invented.
- Whether the result remains recognisably plain text without unintended changes to the content.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Data Ingestion
The knowledge in the files you hand over actually becomes reachable — and the service tells you when it didn't. Graded with the retrieval side held trivially easy (a query in the source's own wording over a small corpus), so a failure is attributable to ingestion rather than to search.
Text Fidelity
Ordinary digital textual content survives completely and correctly — the floor everything downstream stands on; its failure mode reads fine and is invisible
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
Managed RAG
Which managed RAG service turns a folder of private documents into a query API that finds the right passage, answers only from what it retrieved, cites the source that actually supports the claim — and keeps up when the source changes?
Converting a complex PDF into clean Markdown with a hosted API
Also uses this scenario.
Converting a complex PDF into clean Markdown with an open-source library
Also uses this scenario.