Produces directly copyable JSON from the workflow, so the bank-statement extraction is immediately usable without a transformation layer after configuration.
What was measured
Structural Clean Output
Is the JSON directly consumable by a downstream AI pipeline or system without requiring a structural transformation layer?
context, not decisivetransformation
Being directly consumable by a downstream pipeline is valuable, but it is more about integration convenience than whether the tool actually extracts the data correctly. (3 of 3 judges)
What was given, what came back
Test input: Bank Statement PDF · pdf · group: business-document-extraction
Input — what we sent

Research media bank statement 2 jul.png
Bank Statement PDF
A four-page bank statement PDF with dense transaction tables, balance-forward bridges, account metadata, rewards data, and disclaimer text. It was used to stress schema-driven extraction, multi-page continuity, row completeness, and financial numerical accuracy.
Why this input is hard
- · Table extraction across 50+ transaction rows
- · Multi-page continuity with balance-forward bridges
- · Parsing structured account metadata alongside unstructured transaction descriptions
- · Numerical accuracy for deposits, withdrawals, running balances, and summaries
- · Extraction of nested rewards and disclaimer sections
Output — unretouched
Loading file...
Also checked on this input — same tool, 6 other criteria
Extraction Accuracy✓ WorkedDocsumo extracts the bank statement into structured customer, branch, and summary fields, including the account holder name, address, totals, and closing balance, and also supports QA over the document.Extraction Accuracy✓ WorkedExtracts bank-statement account and balance fields into correctly typed values, including account holder 'MR SEENIVASAN', account number '42710540422', opening_balance 114453.65, and closing_balance 116149.46.Extraction Accuracy✓ WorkedCaptures both long bank-statement disclaimer strings as dedicated fields, preserving the insurance_coverage and reporting_period text instead of dropping or flattening it.Schema Adherence✓ WorkedReconstructs the supplied bank-statement schema into a nested JSON object with separate statement.metadata, account_holder, account, branch, balances, transactions, summary, rewards, and disclaimers sections instead of flattening the document into OCR text.Semantic Field Enrichment✓ WorkedDerives transaction_type and transaction_id from the transaction narration, labeling the sample record as UPI with transaction_id 917615251879 instead of leaving only raw description text.Table & Record Completeness◐ MixedOvercounts the bank-statement transaction table in the summary, reporting total_transactions 43 when the researcher says 40 should be counted after excluding balance-forward, tax, and charge entries.
Provenance
- Observation
- 527704b6-b5a4-4c3d-ba5e-e7d4275464b9
- Evidence run
- a061b9e7-a9c5-443d-a171-b296aaf51b8c
- Study
- Extract and query structured data from documents using natural language
- Research task
- 86b9y25e5
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "retab",
scenario: "business-document-extraction"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Structural Clean Output
Extend AI✗ FailedThe output reorders top-level schema objects instead of preserving the declared sequence, so consumers that depend on the original order need an extra transformation step.Landing AI✓ WorkedDelivers JSON that is directly usable downstream, with the demo moving from upload to extraction results and the report noting downloadable output with no extra transformation step.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com