Schema-driven extract.run fully recovers the broken Business Results table, separating the 2024 value and percent-change fields that parse.run had merged together.
What was measured
Advanced Features
Provides separate table/chart extraction and flags low-confidence OCR or ambiguous regions.
context, not decisivetransformation
Separate extraction modes and OCR confidence flags are useful workflow features, but they do not by themselves determine whether the Markdown conversion is good. (3 of 3 judges)
What was given, what came back
Test input: Financial Report - Table Heavy · pdf · group: financial-report-table-heavy
Input — what we sent
A table-heavy corporate financial report used to test extraction of dense, hierarchical financial statements with grouped columns, multi-row headers, segment-reporting tables, and narrative disclosures.
Why this input is hard
- · Multi-page financial statement extraction
- · Hierarchical table reconstruction
- · Grouped columns and multi-row headers
- · Reading order in a report with mixed narrative and tables
- · Document structure retention
- · Markdown usability
Output — unretouched

Also checked on this input — same tool, 7 other criteria
Complex Document Handling◐ MixedProcesses the 18-page report in 9.7 seconds with no truncation; the 9-column segment table at page 17 is flawless even though the first table on page 2 is broken.Markdown Quality✓ WorkedKeeps the syntax clean: no <signature>, <empty>, <b>, <i>, or <u> tags appear anywhere in the output, and the table-of-contents extract renders as a valid pipe table.Reading Order & Structure✗ FailedProvides zero markdown heading markup across all 18 pages; the Roman-numeral sections and numbered subsections are flattened into plain text despite the source’s visible section hierarchy.Table Preservation✓ WorkedPreserves the balance-sheet hierarchy and values, including Assets > Current assets > Cash and deposits, and also reconstructs the later 9-column segment table with all values correct.Table Preservation✗ FailedBreaks a 5-column business-results table by duplicating the 2025 header and merging the 2024 value and percent change into one cell; the Net sales row ends up with 254,811 and 2.6 smashed together.Text & OCR Completeness✓ WorkedConverts all 18 pages with no omitted text; a dense financial-condition paragraph preserves multiple JPY-billion values exactly, including JPY1,252.7 billion, JPY623.6 billion, JPY27.2 billion, JPY10.8 billion, JPY5.0 billion, and JPY9.7 billion.Visual Content Retention✓ WorkedRetains the recurring page-header logo consistently across all 17 pages that carry it; the returned crops on pages 2 and 18 are the same mark and wordmark, with only a 1px width difference (453x36 vs 452x36) consistent with rounding.
Provenance
- Observation
- 991345dd-5c2a-4f09-8dc3-76918b2843d9
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "reducto",
scenario: "financial-report-table-heavy"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on Advanced Features
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com
