Evidence · first-party tested/Best AI Tools for Extracting Structured Data from PDFs and Business Documents
The bank output is rebuilt as nested JSON rather than raw OCR, with branch, account, rewards, balances, summary, and metadata objects populated under the requested statement root.
What was measured
Schema Adherence
Does the output follow the supplied JSON schema hierarchy exactly, with correct nesting, field names, and data types?
decisive for this rankingtransformation
If the output does not match the requested JSON schema exactly, the extracted data cannot be reliably consumed or queried, so this is core to the task. (3 of 3 judges)
What was given, what came back
Test input: Bank Statement PDF · pdf · group: financial-document-extraction
Input — what we sent

Research media bank statement 2 jul.png
Bank Statement PDF
A 4-page bank statement PDF with 51 transactions, balances, rewards, and disclaimer text, used to test schema-driven extraction of dense financial tables and multi-page continuity.
Why this input is hard
- · Table extraction across 50+ transaction rows
- · Multi-page continuity with BALANCE FORWARD bridges
- · Structured metadata vs. free-text transaction descriptions
- · Numerical accuracy for balances, deposits, withdrawals, and summaries
- · Nested schema population for account, branch, balances, rewards, and disclaimers
Also checked on this input — same tool, 5 other criteria
Extraction Accuracy⚠ StruggledIts bank summary aggregation is off: the report says `summary.total_transactions` is 49, while the expected count is 40 after excluding Balance Forward, tax, and charge entries.Semantic Field Enrichment✗ FailedIt leaves derived `transaction_id` values as `null` even when reference identifiers are present in the description, so identifier extraction does not generalize.Semantic Field Enrichment✓ WorkedIt classifies bank transactions into derived `transaction_type` values such as Withdrawal or Deposit from the description text.Structural Clean Output✗ FailedIt does not preserve the authored top-level schema sequence, so consumers expecting the original field order need a transformation step.Table & Record Completeness✓ WorkedThe extractor keeps transaction rows as separate records, and the report says it captured all 51 transactions without merging adjacent rows.
Provenance
- Observation
- a72ab809-ef71-47eb-8461-d9518c96d28b
- Evidence run
- ec4d736d-95f9-4c88-884c-e280435f7b7b
- Study
- Extract and query structured data from documents using natural language
- Research task
- 86b9y25e5
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "extend-ai",
scenario: "financial-document-extraction"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Schema Adherence
Datalab✓ WorkedMaps the bank statement into the requested nested JSON hierarchy instead of flattening it into OCR text, and preserves field-level citation metadata on the extracted objects.Landing AI✓ WorkedReconstructs the bank statement into the requested nested JSON hierarchy, with distinct statement.metadata, account_holder.address, account, branch, statement_period, and balances objects rather than flat OCR text.LlamaParse✓ WorkedKeeps a nested statement schema intact, emitting separate metadata, account_holder, account, branch, statement_period, balances, transactions, summary, rewards, and disclaimers objects instead of flattening the document.Nanonets✓ WorkedIt preserves the requested nested schema directly in the output, populating structured objects such as statement, account, balances, transactions, summary, rewards, and disclaimers instead of flattening the document into OCR text.Reducto✓ WorkedReconstructs the bank statement into a nested JSON structure aligned to the requested schema, with document metadata, account, branch, statement period, transactions, summary, and rewards-style sections instead of flat OCR text.Retab✓ WorkedReconstructs dense statement OCR into the requested nested JSON hierarchy, populating separate statement, account_holder, account, branch, balances, transactions, rewards, and disclaimers objects instead of flattening everything into text.Unstract✓ WorkedPreserves the requested nested statement hierarchy instead of flattening it, with separate metadata, account_holder, account, branch, balances, transactions, summary, and other top-level sections.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com

