Converting a complex PDF into clean Markdown with a hosted API
Evaluates whether a hosted API turns complex PDFs into faithful, usable Markdown without dropping, reordering, or inventing content.
What this benchmark is
This benchmark helps buyers choose a hosted API for PDF-to-Markdown conversion. It asks whether the tool produces Markdown that is good enough for search, review, retrieval, agents, and other downstream systems.
A reader learns whether the API preserves the document's text, order, structure, tables, figures, scanned pages, equations, and code, and whether it handles failures honestly when the source cannot be represented cleanly.
A fictional document set stands in for real customer material so the benchmark can test conversion without exposing private files; the public setup uses source PDFs and checks the resulting Markdown against the original document content and structure.
In scope
- Conversion of complex, real-world PDFs into faithful Markdown.
- Preservation of text, reading order, headings, section structure, tables, figures, charts, scanned text, equations, and code.
- Honest fallback when a source element cannot be expressed cleanly in Markdown.
Out of scope
- Structured field extraction.
- Web-page conversion.
- Form filling.
- Handwriting.
- Document classification and routing.
- PDF creation or editing.
Capabilities included
The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.
| Capability | What it means here | Scenarios |
|---|---|---|
| Text Fidelity | Ordinary digital text survives completely and correctly in the Markdown output. | 1 |
| Reading Order & Layout | The source reading sequence survives the spatial-to-linear conversion without the columns being interleaved or reordered. | 1 |
| Heading & Section Structure | Document structure survives as real Markdown headings and sections, with nesting and footnotes kept in the right relationship. | 2 |
| Table Extraction | Rows, columns, headers, and values survive as a table with the right relationships intact. | 2 |
| Figures & Charts | Visual content survives with its position, caption or title, and a text-reachable trace for later use. | 2 |
| Scanned Document OCR | Text that exists only as pixels is recovered into usable output instead of being silently lost. | 2 |
| Equations & Mathematical Notation | Mathematical notation survives as math markup or an honest fallback, without being turned into misleading prose. | 1 |
| Code Extraction | Code survives as a preformatted fenced block with its whitespace and line structure intact. | 1 |
Scenarios included
A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.
Heading & Section Structure
2 scenariosTable Extraction
2 scenariosScanned Document OCR
2 scenariosHow the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
One test case, one tool
One test case is scored for one tool at a time.
Pass, fail, or partial
Each run ends pass, fail, or partial. A missing capability scores 0 and stays in the arithmetic — a product must not rank higher by having less product. not_applicable, not_measured and blocked (with the reason, never a 0) are distinct states.
Versioned scenarios and resources
Use versioned scenarios and registered resource versions so runs stay comparable.
Recorded stimulus and complete output
Record the input stimulus, the complete output, and the relevant registered references.
Active test cases are partly withheld to reduce benchmark gaming; the public page still lists the capabilities, scenarios, method, and resource types.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
PDF→Markdown fixture corpus — round 1
The registered PDF fixture corpus supplies the versioned inputs used to exercise the benchmark scenarios.