Evidence · first-party tested/Best AI Tools for Extracting Structured Data from PDFs and Business Documents
It classifies bank transactions into derived `transaction_type` values such as Withdrawal or Deposit from the description text.
What was measured
Semantic Field Enrichment
Are derived fields — transaction_type, transaction_id, cheque_number, day patterns, ad codes — correctly classified or extracted beyond raw OCR?
decisive for this rankingtransformation
This ranking is not just about copying OCR text; it also depends on whether the tool can correctly infer or classify document-specific fields needed for useful structured output. (3 of 3 judges)
What was given, what came back
Test input: Bank Statement PDF · pdf · group: financial-document-extraction
Input — what we sent
A 4-page bank statement PDF with 51 transactions, balances, rewards, and disclaimer text, used to test schema-driven extraction of dense financial tables and multi-page continuity.
Why this input is hard
- · Table extraction across 50+ transaction rows
- · Multi-page continuity with BALANCE FORWARD bridges
- · Structured metadata vs. free-text transaction descriptions
- · Numerical accuracy for balances, deposits, withdrawals, and summaries
- · Nested schema population for account, branch, balances, rewards, and disclaimers
Output — unretouched

Also checked on this input — same tool, 4 other criteria
Extraction Accuracy⚠ StruggledIts bank summary aggregation is off: the report says `summary.total_transactions` is 49, while the expected count is 40 after excluding Balance Forward, tax, and charge entries.Schema Adherence✓ WorkedThe bank output is rebuilt as nested JSON rather than raw OCR, with branch, account, rewards, balances, summary, and metadata objects populated under the requested statement root.Structural Clean Output✗ FailedIt does not preserve the authored top-level schema sequence, so consumers expecting the original field order need a transformation step.Table & Record Completeness✓ WorkedThe extractor keeps transaction rows as separate records, and the report says it captured all 51 transactions without merging adjacent rows.
Provenance
- Observation
- 2c4cdb0a-b344-4604-b614-c97af6d74ea2
- Evidence run
- ec4d736d-95f9-4c88-884c-e280435f7b7b
- Study
- Extract and query structured data from documents using natural language
- Research task
- 86b9y25e5
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "extend-ai",
scenario: "financial-document-extraction"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Semantic Field Enrichment
Datalab◐ MixedTransaction typing is inconsistent on merged rows, with one 28 Jun entry labeled Deposit despite showing a 399 withdrawal amount and another 19 Jun merged row labeled Deposit/Withdrawal.Landing AI✗ FailedDoes not preserve source-derived transaction identifiers; it normalizes them into sequential transaction_id values such as "2", "3", and "4", so the embedded IDs in the descriptions cannot be traced back directly.LlamaParse✗ FailedFails to derive transaction-level fields, leaving transaction_id and transaction_type empty even for descriptions that encode ATM, UPI, and CRADJ cues.Nanonets✗ FailedIt does not populate the derived transaction_type field, leaving it null across the statement instead of classifying deposits and withdrawals.Reducto✗ FailedLeaves derived transaction metadata incomplete, with transaction_type and transaction_id missing across transaction rows.Retab✓ WorkedDerives transaction_type and transaction_id on transaction rows, classifying one record as UPI with transaction_id 917615251879 and cheque_number left empty.Unstract✓ WorkedClassifies transaction_type correctly across the transaction array, using Deposit and Withdrawal labels rather than raw OCR text.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com