A cleanly scanned document
A document is scanned and image-only, so its visible text must be recovered without relying on a text layer.
What this scenario means
This scenario checks whether a system can handle scanned pages when the page image is the only source of text. A good system reads the visible content accurately, does not silently treat the page as empty, and does not invent text when the scan has no machine-readable layer. It is a baseline for scanned input, because failure here can otherwise look like a normal document with no content.
What we evaluate
- Whether the visible text from the scanned page is recovered accurately.
- Whether the page is not silently skipped or treated as empty because no text layer exists.
- Whether the output stays faithful to the visible content without inventing text.
- Whether scanned pages are handled as real input, not hidden by omission.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Data Ingestion
The knowledge in the files you hand over actually becomes reachable — and the service tells you when it didn't. Graded with the retrieval side held trivially easy (a query in the source's own wording over a small corpus), so a failure is attributable to ingestion rather than to search.
OCR
The same schema returns the same values when the document is an image, and a mixed document is handled page by page.
Quiz Generation
Creates valid quizzes from supplied material that test the intended knowledge and include sufficient information for answers to be evaluated
Scanned Document OCR
Text that exists only as pixels — a text layer the parser must create; total, silent failure for a whole document class if absent
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
Quiz Generation
Also uses this scenario.
Managed RAG
Which managed RAG service turns a folder of private documents into a query API that finds the right passage, answers only from what it retrieved, cites the source that actually supports the claim — and keeps up when the source changes?
Converting a complex PDF into clean Markdown with a hosted API
Also uses this scenario.
Structured Document Extraction
Also uses this scenario.
Converting a complex PDF into clean Markdown with an open-source library
Also uses this scenario.