Scanned Document OCR
Text that exists only as pixels — a text layer the parser must create; total, silent failure for a whole document class if absent
What this capability means
A tool has this capability when it can recover text that exists only as pixels, by creating an OCR text layer from scanned or image-only PDF content so the words become available in the output. The buyer expects visible text to be read accurately, including scanned pages inside a mixed document, because without OCR that content is a silent total loss.
Boundary: This capability does not judge broader document structure, such as reading order, headings, tables, figures, equations, or code; those are separate capabilities.
Scenarios that test this capability
A scenario is a real-world situation used to test a capability.
A cleanly scanned document
Scenario
Benchmarks that include this capability
A capability has global identity and may be used by more than one benchmark.
Converting a complex PDF into clean Markdown with an open-source library
Also evaluates this capability.
Converting a complex PDF into clean Markdown with a hosted API
Which hosted API converts a complex, real-world PDF into faithful, usable Markdown?