Breaks multirow header relationships in a scanned table, so grouped headers and header-level structure are not reliably preserved.
What was measured
Table Preservation
Preserves complex table structures, including rows, columns, multi-row headers, and merged-cell relationships in markdown.
decisive for this rankingtransformation
Accurate Markdown conversion of complex PDFs depends on keeping table structure intact, not flattening it into plain text. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 4 other criteria
Complex Document Handling✓ WorkedHandles the scanned research paper end-to-end, including multi-column prose, tables, charts, and handwritten marginalia, while still producing parsed markdown.Reading Order & Structure✓ WorkedReconstructs a multi-column research page so the 'STUDY AREA' heading, its paragraphs, and the following 'STAND PRESCRIPTIONS' section stay in order.Text & OCR Completeness◐ MixedDetects faint handwritten margin text, but only partially; the transcription shows 'USDA Semaine' and the remainder is treated as illegible.Visual Content Retention◐ MixedExtracts chart values into a captioned figure block, but the report says the mortality chart's trend visualization is not fully retained.
Provenance
- Observation
- 1e354a89-f9a5-4765-af9a-4cbec79fda3e
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "extend-ai",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Table Preservation
Adobe API✗ FailedBreaks a grouped-column table when intervening text appears, fragmenting the 1979–1981 layout and corrupting the extracted alignment.Landing AI◐ MixedPartially reflects the text that sits between table columns, but the intervening content still disrupts the table structure in the parsed output.LlamaParse✓ WorkedPreserves a nested treatment table with the 7-inch, 10-inch, 12-inch, 100-leave-tree, and clearcut columns and the acres/live-lodgepole rows.Mistral AI✓ WorkedThe multicolumn table is reconstructed without losing its overall layout logic, so the table structure remains readable in the parsed output.Nutrient.io✓ WorkedLargely preserves grouped-column tables, keeping their internal organization intact in the extracted output.Reducto◐ MixedPartially reconstructs Table 1: most of the roughly 90 numeric values are exact, but literal 0 values in the 12-inch column become blanks, one mean cell picks up stray digits (33.0 830000), and a row-label-only section header is broadcast across all six columns in one instance.Tensorlake✗ FailedStruggles with hierarchical scanned tables, misplacing column headers and producing unreliable reconstructions on both the multicolumn table and the denser complex table.Upstage AI◐ MixedKeeps the table values intact but reconstructs the headers incorrectly, leaving a grouped-column table with inconsistent structure.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
