Provides separate chart extraction for the SG&A waterfall, outputting chart metadata with a title, axis labels, 10 category labels, and a numeric series instead of only inline prose.
What was measured
Advanced Features
Provides separate table/chart extraction and flags low-confidence OCR or ambiguous regions.
context, not decisivetransformation
Separate extraction modes and OCR confidence flags are useful workflow features, but they do not by themselves determine whether the Markdown conversion is good. (3 of 3 judges)
What was given, what came back
Test input: Hybrid Earnings Report · pdf · group: hybrid-earnings-report
Input — what we sent

Hybrid Earnings Report
A long, real-world hybrid annual report used to benchmark end-to-end PDF-to-markdown conversion. It combines native text, financial tables, charts/graphics, and scanned signature/stamp regions, stressing preservation of document structure, reading order, and embedded visual content across a multi-section report.
Why this input is hard
- · Native digital text extraction
- · Complex financial table preservation
- · Chart and graphic retention
- · Image-based signatures and stamps
- · Reading order across a long multi-section report
- · Markdown quality and consistency
Output — unretouched

Also checked on this input — same tool, 6 other criteria
Complex Document Handling✓ WorkedProcesses an 84-page mixed-content annual report without collapsing the hierarchy, keeping text, tables, charts, and scanned signatures usable within the extracted workflow.Markdown Quality✓ WorkedRenders the extraction as structured, copyable markdown in the Document Markdown view rather than a flat text dump.Reading Order & Structure✓ WorkedPreserves heading order and section relationships in an 84-page hybrid annual report, keeping the narrative, bullets, and figure placement aligned with the source flow.Table Preservation✓ WorkedKeeps a 5-year financial summary table structured, retaining the row/column relationships across sales, expenses, EBIT, and per-share rows instead of flattening the table.Text & OCR Completeness◐ MixedOn a blurry Ernst & Young signoff, the OCR preserves the firm reference but makes a symbol-level mistake by rendering the ampersand as '+', so the text is close but not exact.Visual Content Retention✓ WorkedRetains scanned signature-page visuals as figure blocks plus signer text, including named executives and dates, rather than dropping the signature regions from the output.
Provenance
- Observation
- 324c753c-f3ab-4165-8bd1-eda961278d78
- Evidence run
- 6e3160de-fe46-4b45-b071-72560b5c5d0e
- Study
- Convert a Complex PDF into Clean Markdown with an API
- Research task
- 86b9h7t37
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "tensorlake",
scenario: "hybrid-earnings-report"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Advanced Features
Landing AI✓ WorkedEmits semantic attestation blocks for handwritten-signature regions and a blurry signature stamp, distinguishing the signer/company name and the signature legibility instead of leaving those regions unannotated.Reducto✓ WorkedThe separate schema-driven extract.run endpoint cleanly recovers both a financial table and a donut chart: all requested table rows and all five donut percentages come back exactly.Upstage AI✓ WorkedSeparately extracts chart content into textual summaries and value tables, including a waterfall chart and bar charts with year/value pairs and growth statistics.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com