Struggles with hierarchical scanned tables, misplacing column headers and producing unreliable reconstructions on both the multicolumn table and the denser complex table.
What was measured
Table Preservation
Preserves complex table structures, including rows, columns, multi-row headers, and merged-cell relationships in markdown.
decisive for this rankingtransformation
Accurate Markdown conversion of complex PDFs depends on keeping table structure intact, not flattening it into plain text. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent




Scanned Research Paper
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched


Also checked on this input — same tool, 4 other criteria
Advanced Features✓ WorkedPerforms dedicated chart extraction on a scanned bar chart, turning the figure into structured chart content with year-by-year values, treatment labels, and the 'CUT COMPLETED' annotation.Complex Document Handling◐ MixedOn the scanned research paper, section flow and chart extraction work, but hierarchical tables degrade, so mixed-content handling is uneven rather than consistently robust.Markdown Quality✓ WorkedRenders the scanned-paper extraction as structured markdown in the tool workflow, rather than only exposing raw OCR text.Reading Order & Structure✓ WorkedRetains section-level reading order in a scanned multi-column paper, with headings continuing to guide the flow across columns and into the next section.
Provenance
- Observation
- c19621dc-83eb-467f-9210-161a2adcac66
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "tensorlake",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Table Preservation
Adobe API✗ FailedBreaks a grouped-column table when intervening text appears, fragmenting the 1979–1981 layout and corrupting the extracted alignment.Extend AI⚠ StruggledBreaks multirow header relationships in a scanned table, so grouped headers and header-level structure are not reliably preserved.Landing AI✓ WorkedPreserves a complex diameter-class table with before-cut, trees-cut-per-acre, and after-cut relationships across the treatment rows and check area.LlamaParse✓ WorkedPreserves a nested treatment table with the 7-inch, 10-inch, 12-inch, 100-leave-tree, and clearcut columns and the acres/live-lodgepole rows.Mistral AI✓ WorkedThe multicolumn table is reconstructed without losing its overall layout logic, so the table structure remains readable in the parsed output.Nutrient.io✓ WorkedLargely preserves grouped-column tables, keeping their internal organization intact in the extracted output.Reducto◐ MixedPartially reconstructs Table 1: most of the roughly 90 numeric values are exact, but literal 0 values in the 12-inch column become blanks, one mean cell picks up stray digits (33.0 830000), and a row-label-only section header is broadcast across all six columns in one instance.Upstage AI◐ MixedKeeps the table values intact but reconstructs the headers incorrectly, leaving a grouped-column table with inconsistent structure.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com