Recovers readable text from scanned pages, including the abstract, keywords, author affiliations, title, and opening paragraphs on the first page.
What was measured
Text & OCR Completeness
Extracts all readable content, including scanned pages, with accurate OCR and minimal omissions.
decisive for this rankingtransformation
If the tool misses readable text or fails on scanned pages, it has not actually converted the PDF faithfully into Markdown. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched


Also checked on this input — same tool, 6 other criteria
Complex Document Handling✓ WorkedProcesses an image-only scanned research paper end to end and returns a parsed markdown output plus preview, showing it can handle a multi-page scanned document.Reading Order & Structure✗ FailedMisorders the first page so the abstract is placed before the title, showing a layout-reading-order inversion on the scanned article.Reading Order & Structure✓ WorkedThe tool preserves section hierarchy and column order in a dense multi-column scanned section, keeping headings aligned with the correct body text.Table Preservation✓ WorkedLargely preserves grouped-column tables, keeping their internal organization intact in the extracted output.Table Preservation✗ FailedBreaks more complex multi-level tables, with rows and merged cells misaligned or lost once the hierarchy becomes dense.Visual Content Retention⚠ StruggledExtracts the Figure 3 chart's values, but the chart structure and layout are not preserved, so the visualization is flattened into text-like output.
Provenance
- Observation
- 1ff9ad63-6629-4bd0-a643-ff827f73c602
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "nutrient-io",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Text & OCR Completeness
Adobe API✓ WorkedRecovers the visible title, abstract, keywords, and opening paragraphs from a scanned USDA forestry report as dense OCR text.Extend AI◐ MixedDetects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.Landing AI◐ MixedOCRs most of the scanned cover-page text, including the agency header, report number/date, title, authors, abstract, and keywords, but inserts noisy tokens such as '186153' and 'USA/-' into the title line.LlamaParse✓ WorkedRecovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs.Mistral AI◐ MixedThe report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy.PDFVector✓ WorkedSuccessfully parsed a 12-page scanned research paper and produced a long extracted-text preview in 20.6 s using 48 credits.Reducto✓ WorkedConverts all 12 scanned pages with no gaps; the page-11-to-page-12 handoff is preserved verbatim, and a separate page-marker output shows pages 1 through 12 present with no missing markers.Upstage AI✓ WorkedOCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
