The report says the parser recovers much of the underlying text from the scanned paper, but it does not present a measured completeness rate and the first-page hierarchy is still lossy.
What was measured
Text & OCR Completeness
Extracts all readable content, including scanned pages, with accurate OCR and minimal omissions.
decisive for this rankingtransformation
If the tool misses readable text or fails on scanned pages, it has not actually converted the PDF faithfully into Markdown. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 7 other criteria
Complex Document Handling✓ WorkedThe tool processes a scanned multi-column research paper end-to-end and returns OCR text, tables, and embedded chart assets in page-wise markdown output.Markdown Quality✓ WorkedThe export is packaged as usable markdown files in a ZIP, with both overall and page-wise outputs available for inspection.Reading Order & Structure✓ WorkedSection hierarchy and reading flow are preserved in the scanned paper, keeping headings and supporting paragraphs correctly connected despite the multi-column layout.Reading Order & Structure✗ FailedThe opening page loses the distinction between the document title and the abstract, flattening the semantic organization of the first page.Table Preservation✓ WorkedThe multicolumn table is reconstructed without losing its overall layout logic, so the table structure remains readable in the parsed output.Table Preservation✗ FailedBroken column boundaries and disrupted value alignment make the reconstructed table significantly less faithful to the source.Visual Content Retention✓ WorkedCharts are exposed through both page-wise markdown files and the extracted visual assets, keeping the visual content linked to its original document location.
Provenance
- Observation
- 932e3296-1f2a-4d58-9bb4-5c7bb1e269a7
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "mistral-ai",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Text & OCR Completeness
Adobe API✓ WorkedRecovers the visible title, abstract, keywords, and opening paragraphs from a scanned USDA forestry report as dense OCR text.Extend AI◐ MixedDetects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.Landing AI◐ MixedOCRs most of the scanned cover-page text, including the agency header, report number/date, title, authors, abstract, and keywords, but inserts noisy tokens such as '186153' and 'USA/-' into the title line.LlamaParse✓ WorkedRecovers dense readable prose from a scanned page-image source, including the section heading and multiple long paragraphs.Nutrient.io✓ WorkedRecovers readable text from scanned pages, including the abstract, keywords, author affiliations, title, and opening paragraphs on the first page.PDFVector✓ WorkedSuccessfully parsed a 12-page scanned research paper and produced a long extracted-text preview in 20.6 s using 48 credits.Reducto✓ WorkedConverts all 12 scanned pages with no gaps; the page-11-to-page-12 handoff is preserved verbatim, and a separate page-marker output shows pages 1 through 12 present with no missing markers.Upstage AI✓ WorkedOCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
