A document mixing digital and scanned pages
A document contains both digital text pages and scanned image-only pages.
What this scenario means
This is a real test because one page can look like ordinary text while another page is image-only and needs separate handling. A good system detects the scanned page, processes pages one by one, and does not let the mixed format hide or skip that page.
What we evaluate
- Whether the scanned page inside a mixed document is processed rather than silently skipped.
- Whether page-by-page handling keeps content from both digital text pages and scanned image-only pages.
- Whether the system treats the document as mixed, not as a fully digital file with no special handling needed.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
OCR
The same schema returns the same values when the document is an image, and a mixed document is handled page by page.
Scanned Document OCR
Text that exists only as pixels — a text layer the parser must create; total, silent failure for a whole document class if absent
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
Converting a complex PDF into clean Markdown with a hosted API
Which hosted API converts a complex, real-world PDF into faithful, usable Markdown?
Structured Document Extraction
Also uses this scenario.
Converting a complex PDF into clean Markdown with an open-source library
Also uses this scenario.