Separately extracts Figure 3 into year-by-year values for 1972–1981 and tags it as an image asset, showing chart-specific extraction beyond plain OCR.
What was measured
Advanced Features
Provides separate table/chart extraction and flags low-confidence OCR or ambiguous regions.
context, not decisivetransformation
Separate extraction modes and OCR confidence flags are useful workflow features, but they do not by themselves determine whether the Markdown conversion is good. (3 of 3 judges)
What was given, what came back
Test input: Scanned Research Paper · pdf · group: scanned-research-paper
Input — what we sent
An image-only scanned research paper used to stress OCR and layout recovery in a multi-column academic document with figures, charts, tables, captions, and references.
Why this input is hard
- · OCR on scanned pages
- · Multi-column reading order
- · Figure and chart handling
- · Table reconstruction from scans
- · Caption association
- · Reference extraction
- · Overall document structure retention
Output — unretouched

Also checked on this input — same tool, 3 other criteria
Reading Order & Structure✗ FailedBreaks paragraph-level segmentation in the scanned two-column page, so the extracted text no longer follows the source column order cleanly.Table Preservation◐ MixedKeeps the table values intact but reconstructs the headers incorrectly, leaving a grouped-column table with inconsistent structure.Text & OCR Completeness✓ WorkedOCRs dense scanned prose successfully, capturing the ABSTRACT heading and multiple paragraphs of body text rather than only captions or labels.
Provenance
- Observation
- 284ba196-edbd-492d-a9db-504bfe85f361
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "upstage-ai",
scenario: "scanned-research-paper"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Advanced Features
Reducto✓ WorkedSchema-driven extract.run fully recovers the corrupted Table 4 rows, returning all 12 checked rows exactly, including the 10-inch-cut 1979 MPB row whose parse.run output was badly garbled.Tensorlake✓ WorkedPerforms dedicated chart extraction on a scanned bar chart, turning the figure into structured chart content with year-by-year values, treatment labels, and the 'CUT COMPLETED' annotation.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
