Reducto in Converting a complex PDF into clean Markdown with a hosted API
Scenario-level performance from current published Results.
5 scenarios with published Results · 12 scenarios in the benchmark
How Reducto performed
Open a capability to explore its scenarios. Each row reports the test set in its published Result; counts are not combined into an overall score.
Code Extraction1 scenario · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A document containing a code block | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedReducto did not preserve the code block as fenced or preformatted Markdown. The returned file had no fences or indented code lines, while the source block’s line breaks remained and the output also altered code tokens, including `window = reload_window(reset_at, margin_minutes=45)` becoming `window reload_window (reset_at, margin_minutes=45)`. | 1/1 assessed1/1 gradablePublished test set | View Result → |
Equations & Mathematical Notation1 scenario · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A document containing mathematical equations | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedReducto failed a document with mathematical equations: the returned Markdown contained no math markup and no fallback, and the equations were flattened into prose with missing symbols and substituted characters. This is the case: the source page shows set mathematics, while the output preserves only garbled lines instead of LaTeX, MathML, or an explicit fallback. | 1/1 assessed1/1 gradablePublished test set | View Result → |
Heading & Section Structure2 scenarios · 1 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A document with styled headings and subheadings | 0 Pass1 Fail0 Not gradable 1 of 1 test case failedWhat happenedReducto flattened the heading structure in the Markdown output. The file contains one "#", 26 "##" headings, and no deeper levels; the document title is plain paragraph text. The two visible parent/child pairs were emitted as peers, so the nesting expected for styled headings and subheadings was not preserved. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| Footnotes at the bottom of the page | No published result | ||
Scanned Document OCR2 scenarios · 2 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A cleanly scanned document | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedReducto accurately recovered the visible text from a cleanly scanned document. The scan’s text matched the output field by field, with the handwritten signature left as a <signature> placeholder and only two middle-dot separators dropped from the letterhead and footer lines. | 1/1 assessed1/1 gradablePublished test set | View Result → |
| A document mixing digital and scanned pages | 1 Pass0 Fail0 Not gradable 1 of 1 test case passedWhat happenedThe mixed PDF’s scanned page was processed rather than skipped. Reducto’s Markdown includes the image-only page 2 with its headers, table, certificate section, and recipient details, matching the digital pages’ presentation in the checked evidence. | 1/1 assessed1/1 gradablePublished test set | View Result → |
Figures & Charts2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A document with a data chart | No published result | ||
| A document with figures and captions | No published result | ||
Reading Order & Layout1 scenario · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A page laid out in multiple columns | No published result | ||
Table Extraction2 scenarios · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| A simple, clearly formatted table | No published result | ||
| A table that continues across a page break | No published result | ||
Text Fidelity1 scenario · 0 with published Results
Capability in this benchmark →
| Scenario | Published outcomes | Test coverage | Result |
|---|---|---|---|
| An ordinary digital text document | No published result | ||
Reading these Results
Published evidence and test coverage answer different questions.
Which scenarios have a Result?
A published Result is public evidence for this tool on one scenario. “No published result” does not say whether testing has taken place.
What does each Result cover?
Assessed includes Pass, Fail and Not gradable. Gradable includes Pass and Fail. Both use the pinned test count in that published Result.
Inventory is not testing progress
The 12 scenarios describe this benchmark’s scope. They are not an assumed applicability or test-coverage denominator for Reducto.