Evidence · first-party tested/Best AI Tools for Extracting Structured Data from PDFs and Business Documents
It does not populate the derived transaction_type field, leaving it null across the statement instead of classifying deposits and withdrawals.
What was measured
Semantic Field Enrichment
Are derived fields — transaction_type, transaction_id, cheque_number, day patterns, ad codes — correctly classified or extracted beyond raw OCR?
decisive for this rankingtransformation
This ranking is not just about copying OCR text; it also depends on whether the tool can correctly infer or classify document-specific fields needed for useful structured output. (3 of 3 judges)
What was given, what came back
Test input: Bank Statement PDF · pdf · group: financial-document-extraction
Input — what we sent

Research media bank statement 2 jul.png
Bank Statement PDF
A 4-page bank statement PDF with 51 transactions, balances, rewards, and disclaimer text, used to test schema-driven extraction of dense financial tables and multi-page continuity.
Why this input is hard
- · Table extraction across 50+ transaction rows
- · Multi-page continuity with BALANCE FORWARD bridges
- · Structured metadata vs. free-text transaction descriptions
- · Numerical accuracy for balances, deposits, withdrawals, and summaries
- · Nested schema population for account, branch, balances, rewards, and disclaimers
Also checked on this input — same tool, 6 other criteria
Extraction Accuracy✓ WorkedIt accurately fills core statement metadata and balances, including account number 42710540422, opening balance 114453.65, and closing balance 116149.46.Extraction Accuracy◐ MixedIt leaves date and value_date blank on multiple transactions even when the source row contains them; the report says 15+ transactions are affected.Extraction Accuracy✗ FailedIt leaves summary.total_transactions null even though the statement has 51 transactions, so the summary count is not extracted as a usable value.Schema Adherence✓ WorkedIt preserves the requested nested schema directly in the output, populating structured objects such as statement, account, balances, transactions, summary, rewards, and disclaimers instead of flattening the document into OCR text.Table & Record Completeness✗ FailedIt merges adjacent bank-statement rows into a single record description, so one extracted transaction can absorb neighboring content rather than staying row-bounded.Table & Record Completeness✗ FailedIt undercounts the transaction table, extracting 47 transactions when the statement actually contains 51, so 4 records are missing.
Provenance
- Observation
- db96ee62-b038-4189-9112-9ff53de816f8
- Evidence run
- ec4d736d-95f9-4c88-884c-e280435f7b7b
- Study
- Extract and query structured data from documents using natural language
- Research task
- 86b9y25e5
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "nanonets",
scenario: "financial-document-extraction"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Semantic Field Enrichment
Datalab✗ FailedDoes not populate schema-derived transaction identifiers at all: the 18 Jun withdrawal keeps transaction_id = null even though the identifier is visible in the source row.Extend AI✗ FailedIt leaves derived `transaction_id` values as `null` even when reference identifiers are present in the description, so identifier extraction does not generalize.Landing AI✓ WorkedAdds meaningful transaction_type labels to extracted rows, classifying the sample 18 Jun records as Withdrawal, Withdrawal, and Deposit instead of leaving the field as raw OCR text.LlamaParse✗ FailedFails to derive transaction-level fields, leaving transaction_id and transaction_type empty even for descriptions that encode ATM, UPI, and CRADJ cues.Reducto✗ FailedLeaves derived transaction metadata incomplete, with transaction_type and transaction_id missing across transaction rows.Retab✓ WorkedDerives transaction_type and transaction_id on transaction rows, classifying one record as UPI with transaction_id 917615251879 and cheque_number left empty.Unstract✓ WorkedClassifies transaction_type correctly across the transaction array, using Deposit and Withdrawal labels rather than raw OCR text.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com