--- title: "Best AI Tools for Structured Document Extraction from Bank Statements and Invoices" type: "Ranking" url: "https://aidemos.com/best/document-extraction" description: "We tested seven AI document extraction tools on a dense bank statement PDF and a structured broadcast invoice PDF, checking schema adherence, row-level completeness, transaction IDs, line items, totals, confidence metadata, and export-ready JSON." readTime: "14 min read" tested: "Landing AI vs Retab vs Extend AI vs LlamaParse vs Datalab vs Reducto vs Nanonets" testedDate: "June 2026" category: "developer-tools" published: "2026-07-01T10:51:09.777700+00:00" updated: "2026-07-24T07:06:24.748042+00:00" evidenceCount: 128 verifiedCount: 0 coverage: "partial" --- # Best AI Tools for Structured Document Extraction from Bank Statements and Invoices `7 Tools Tested` · `Bank Statement PDF` · `Invoice PDF` · `51 Transactions` · `8 Line Items` · `Confidence Scores` **Tested:** Landing AI vs Retab vs Extend AI vs LlamaParse vs Datalab vs Reducto vs Nanonets · June 2026 > We tested seven AI document extraction tools on a dense bank statement PDF and a structured broadcast invoice PDF, checking schema adherence, row-level completeness, transaction IDs, line items, totals, confidence metadata, and export-ready JSON. ## Our Verdict **#1 pick: Landing AI** (Best) — Strong schema adherence and complete row/line-item extraction, but it invents sequential transaction IDs and occasionally misreads alphanumeric codes. - #2 Retab — Strong at nested schema reconstruction and semantic field labeling, with a few cleanup issues in record counts and final formatting. - #3 Extend AI — Strong structured extraction with a few cleanup gaps in ordering and field normalization. - #4 LlamaParse — Strong at producing clean, schema-shaped JSON, but it slips on record counts and derived transaction fields. - #5 Datalab — Strong on schema-shaped invoice output and cited metadata, but bank-statement transaction handling breaks down on row boundaries and derived transaction fields. - #6 Reducto — Strong at reconstructing invoices and structured JSON, but the bank statement shows reliability gaps in reconciliation fields and derived transaction metadata. - #7 Nanonets — Strong schema and export handling, but weaker on bank-statement accuracy and derived-field enrichment. ## How We Tested We uploaded the same two real-world documents to every tool: a multi-page bank statement PDF and a two-page broadcast invoice PDF. Each tool was tested against the matching custom JSON schema, and we compared schema adherence, field accuracy, transaction and line-item completeness, semantic enrichment such as transaction IDs and transaction types, and whether the final JSON was directly consumable without post-processing. Cycle 1 focused on structured extraction only; natural-language querying, validation UI behavior, and scanned/low-resolution stress inputs were not part of this round. **What we evaluated:** | Criterion | Description | | --- | --- | | Schema Adherence | Does the output follow the supplied JSON schema hierarchy exactly, with correct nesting, field names, and data types? | | Extraction Accuracy | Are field values correct, complete, and free of OCR or parsing errors, including numerical precision on financial fields? | | Table & Record Completeness | Are all tabular rows or records extracted without merging, duplication, omission, or phantom records? | | Semantic Field Enrichment | Are derived fields such as transaction_type, transaction_id, cheque_number, day patterns, and ad codes correctly classified or extracted beyond raw OCR? | | Structural Clean Output | Is the JSON directly consumable by a downstream AI pipeline or system without requiring a structural transformation layer? | ## Evidence (first-party, tested) *128 tested cells · 0/128 artifact-verified. Scores are out of 5. Cite a cell by its Evidence ID, e.g. `ev:datalab·bank-statement-pdf·extraction-accuracy`.* | Tool | Criterion | Scenario | Verdict | Score | Tested | Proof | Evidence ID | | --- | --- | --- | --- | --- | --- | --- | --- | | Datalab | Extraction Accuracy | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-datalab-bank-statement-extracted-details-9b3fa938f6ae.png) | `ev:datalab·bank-statement-pdf·extraction-accuracy` | | Datalab | Extraction Accuracy | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-datalab-invoice-summary-485f182a6662.png) | `ev:datalab·invoice-pdf·extraction-accuracy` | | Datalab | Schema Adherence | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-datalab-bank-statement-extracted-details-9b3fa938f6ae.png) | `ev:datalab·cross·schema-adherence` | | Datalab | Semantic Field Enrichment | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:datalab·invoice-pdf·semantic-field-enrichment` | | Datalab | Semantic Field Enrichment | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-bank-statement-18-june-transaction-id-a47114d15795.png) | `ev:datalab·bank-statement-pdf·semantic-field-enrichment` | | Datalab | Structural Clean Output | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-datalab-bank-statement-extracted-citatio-85bb110cea50.png) | `ev:datalab·cross·structural-clean-output` | | Datalab | Structural Clean Output | Invoice PDF | ◐ mixed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-invoice-schema-order-1938d10e3382.png) | `ev:datalab·invoice-pdf·structural-clean-output` | | Datalab | Table & Record Completeness | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-bank-statement-18-june-transaction-id-a47114d15795.png) | `ev:datalab·bank-statement-pdf·table-and-record-completeness` | | Datalab | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:datalab·invoice-pdf·table-and-record-completeness` | | Extend AI | Advanced Features (Bonus) | cross-scenario | ✓ worked | — | — | — | `ev:extend-ai·cross·advanced-features-bonus` | | Extend AI | Advanced Features (Bonus) | Scanned Research Paper INT 1983-07 Issue 333 | ✓ worked | — | — | — | `ev:extend-ai·scanned-research-paper-int-1983-07-issue-333·advanced-features-bonus` | | Extend AI | Complex Document Handling | cross-scenario | ✓ worked | — | — | — | `ev:extend-ai·cross·complex-document-handling` | | Extend AI | Extraction Accuracy | Invoice PDF | ◐ mixed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-invoice-summary-2-6506753307a7.png) | `ev:extend-ai·invoice-pdf·extraction-accuracy` | | Extend AI | Extraction Accuracy | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-bank-statement-summary-a4a38bf7978f.png) | `ev:extend-ai·bank-statement-pdf·extraction-accuracy` | | Extend AI | Markdown Quality | cross-scenario | ✓ worked | — | — | — | `ev:extend-ai·cross·markdown-quality` | | Extend AI | Reading Order & Structure | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:extend-ai·hybrid-earnings-report·reading-order-structure` | | Extend AI | Reading Order & Structure | Financial Report - Table Heavy | ✓ worked | — | — | — | `ev:extend-ai·financial-report-table-heavy·reading-order-structure` | | Extend AI | Reading Order & Structure | Scanned Research Paper | ✓ worked | — | — | — | `ev:extend-ai·scanned-research-paper·reading-order-structure` | | Extend AI | Reading Order & Structure | Financial Report | ✓ worked | — | — | — | `ev:extend-ai·financial-report·reading-order-structure` | | Extend AI | Reading Order & Structure | Scanned Research Paper INT 1983-07 Issue 333 | ✓ worked | — | — | — | `ev:extend-ai·scanned-research-paper-int-1983-07-issue-333·reading-order-structure` | | Extend AI | Reading Order & Structure | Sumitomo Heavy Industries Consolidated Financial Report | ✓ worked | — | — | — | `ev:extend-ai·sumitomo-heavy-industries-consolidated-financial-report·reading-order-structure` | | Extend AI | Schema Adherence | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-document-level-bank-statement-6a3ae27f2afa.png) | `ev:extend-ai·bank-statement-pdf·schema-adherence` | | Extend AI | Schema Adherence | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-invoice-schema-order-1938d10e3382.png) | `ev:extend-ai·invoice-pdf·schema-adherence` | | Extend AI | Semantic Field Enrichment | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-extracted-transaction-id-fc2e1cc3385b.png) | `ev:extend-ai·bank-statement-pdf·semantic-field-enrichment` | | Extend AI | Semantic Field Enrichment | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:extend-ai·invoice-pdf·semantic-field-enrichment` | | Extend AI | Structural Clean Output | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-invoice-extracted-metadata-11343985f169.png) | `ev:extend-ai·cross·structural-clean-output` | | Extend AI | Structural Clean Output | Invoice PDF | ◐ mixed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-invoice-schema-order-1938d10e3382.png) | `ev:extend-ai·invoice-pdf·structural-clean-output` | | Extend AI | Structural Clean Output | Bank Statement PDF | ◐ mixed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-bank-statement-schema-order-69fcdddefcea.png) | `ev:extend-ai·bank-statement-pdf·structural-clean-output` | | Extend AI | Table & Record Completeness | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/extend-ai-extend-ai-extracted-transaction-item-4618ec1646be.png) | `ev:extend-ai·bank-statement-pdf·table-and-record-completeness` | | Extend AI | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:extend-ai·invoice-pdf·table-and-record-completeness` | | Extend AI | Table Preservation | Scanned Research Paper | ✓ worked | — | — | — | `ev:extend-ai·scanned-research-paper·table-preservation` | | Extend AI | Table Preservation | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:extend-ai·hybrid-earnings-report·table-preservation` | | Extend AI | Table Preservation | Financial Report - Table Heavy | ✓ worked | — | — | — | `ev:extend-ai·financial-report-table-heavy·table-preservation` | | Extend AI | Table Preservation | Financial Report | ✗ failed | — | — | — | `ev:extend-ai·financial-report·table-preservation` | | Extend AI | Table Preservation | Scanned Research Paper INT 1983-07 Issue 333 | ◐ mixed | — | — | — | `ev:extend-ai·scanned-research-paper-int-1983-07-issue-333·table-preservation` | | Extend AI | Table Preservation | Sumitomo Heavy Industries Consolidated Financial Report | ✗ failed | — | — | — | `ev:extend-ai·sumitomo-heavy-industries-consolidated-financial-report·table-preservation` | | Extend AI | Text & OCR Completeness | Scanned Research Paper | ✗ failed | — | — | — | `ev:extend-ai·scanned-research-paper·text-ocr-completeness` | | Extend AI | Text & OCR Completeness | Hybrid Earnings Report | ◐ mixed | — | — | — | `ev:extend-ai·hybrid-earnings-report·text-ocr-completeness` | | Extend AI | Visual Content Retention | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:extend-ai·hybrid-earnings-report·visual-content-retention` | | Extend AI | Visual Content Retention | Scanned Research Paper | ⚠ struggled | — | — | — | `ev:extend-ai·scanned-research-paper·visual-content-retention` | | Extend AI | Visual Content Retention | Scanned Research Paper INT 1983-07 Issue 333 | ◐ mixed | — | — | — | `ev:extend-ai·scanned-research-paper-int-1983-07-issue-333·visual-content-retention` | | Extend AI | Visual Content Retention | Target 2015 Annual Report | ⚠ struggled | — | — | — | `ev:extend-ai·target-2015-annual-report·visual-content-retention` | | Landing AI | Advanced Features (Bonus) | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:landing-ai·hybrid-earnings-report·advanced-features-bonus` | | Landing AI | Advanced Features (Bonus) | Target 2015 Annual Report | ✓ worked | — | — | — | `ev:landing-ai·target-2015-annual-report·advanced-features-bonus` | | Landing AI | Complex Document Handling | cross-scenario | ✓ worked | — | — | — | `ev:landing-ai·cross·complex-document-handling` | | Landing AI | Extraction Accuracy | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/landing-ai-landing-ai-extracted-bank-statement-c4910d90b9f3.png) | `ev:landing-ai·bank-statement-pdf·extraction-accuracy` | | Landing AI | Extraction Accuracy | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/landing-ai-landing-ai-invoice-summary-eb031f81c889.png) | `ev:landing-ai·invoice-pdf·extraction-accuracy` | | Landing AI | Markdown Quality | cross-scenario | ✓ worked | — | — | — | `ev:landing-ai·cross·markdown-quality` | | Landing AI | Reading Order & Structure | Hybrid Earnings Report | ✗ failed | — | — | — | `ev:landing-ai·hybrid-earnings-report·reading-order-structure` | | Landing AI | Reading Order & Structure | Financial Report - Table Heavy | ⚠ struggled | — | — | — | `ev:landing-ai·financial-report-table-heavy·reading-order-structure` | | Landing AI | Reading Order & Structure | Scanned Research Paper | ✓ worked | — | — | — | `ev:landing-ai·scanned-research-paper·reading-order-structure` | | Landing AI | Reading Order & Structure | Financial Report | ✗ failed | — | — | — | `ev:landing-ai·financial-report·reading-order-structure` | | Landing AI | Reading Order & Structure | Sumitomo Heavy Industries Consolidated Financial Report | ✓ worked | — | — | — | `ev:landing-ai·sumitomo-heavy-industries-consolidated-financial-report·reading-order-structure` | | Landing AI | Reading Order & Structure | Target 2015 Annual Report | ✗ failed | — | — | — | `ev:landing-ai·target-2015-annual-report·reading-order-structure` | | Landing AI | Schema Adherence | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/landing-ai-landing-ai-extracted-invoice-metadata-57a291c39e3d.png) | `ev:landing-ai·invoice-pdf·schema-adherence` | | Landing AI | Schema Adherence | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/landing-ai-landing-ai-extracted-bank-statement-c4910d90b9f3.png) | `ev:landing-ai·bank-statement-pdf·schema-adherence` | | Landing AI | Semantic Field Enrichment | Invoice PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:landing-ai·invoice-pdf·semantic-field-enrichment` | | Landing AI | Semantic Field Enrichment | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/landing-ai-landing-ai-bank-statement-18-jun-extract-321aa2936cc6.png) | `ev:landing-ai·bank-statement-pdf·semantic-field-enrichment` | | Landing AI | Structural Clean Output | cross-scenario | ✓ worked | — | — | — | `ev:landing-ai·cross·structural-clean-output` | | Landing AI | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:landing-ai·invoice-pdf·table-and-record-completeness` | | Landing AI | Table & Record Completeness | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-bank-statement-21-jun-transactions-890dc245679d.png) | `ev:landing-ai·bank-statement-pdf·table-and-record-completeness` | | Landing AI | Table Preservation | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:landing-ai·hybrid-earnings-report·table-preservation` | | Landing AI | Table Preservation | Financial Report - Table Heavy | ✗ failed | — | — | — | `ev:landing-ai·financial-report-table-heavy·table-preservation` | | Landing AI | Table Preservation | Scanned Research Paper | ⚠ struggled | — | — | — | `ev:landing-ai·scanned-research-paper·table-preservation` | | Landing AI | Table Preservation | Financial Report | ⚠ struggled | — | — | — | `ev:landing-ai·financial-report·table-preservation` | | Landing AI | Table Preservation | Sumitomo Heavy Industries Consolidated Financial Report | ⚠ struggled | — | — | — | `ev:landing-ai·sumitomo-heavy-industries-consolidated-financial-report·table-preservation` | | Landing AI | Table Preservation | Target 2015 Annual Report | ✓ worked | — | — | — | `ev:landing-ai·target-2015-annual-report·table-preservation` | | Landing AI | Text & OCR Completeness | Scanned Research Paper | ✓ worked | — | — | — | `ev:landing-ai·scanned-research-paper·text-ocr-completeness` | | Landing AI | Visual Content Retention | Hybrid Earnings Report | ✗ failed | — | — | — | `ev:landing-ai·hybrid-earnings-report·visual-content-retention` | | Landing AI | Visual Content Retention | Scanned Research Paper | ✗ failed | — | — | — | `ev:landing-ai·scanned-research-paper·visual-content-retention` | | Landing AI | Visual Content Retention | Target 2015 Annual Report | ✗ failed | — | — | — | `ev:landing-ai·target-2015-annual-report·visual-content-retention` | | LlamaParse | Advanced Features (Bonus) | Scanned Research Paper | ✓ worked | — | — | — | `ev:llamaparse·scanned-research-paper·advanced-features-bonus` | | LlamaParse | Advanced Features (Bonus) | cross-scenario | ✓ worked | — | — | — | `ev:llamaparse·cross·advanced-features-bonus` | | LlamaParse | Complex Document Handling | cross-scenario | ✓ worked | — | — | — | `ev:llamaparse·cross·complex-document-handling` | | LlamaParse | Complex Document Handling | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:llamaparse·hybrid-earnings-report·complex-document-handling` | | LlamaParse | Complex Document Handling | Target 2015 Annual Report | ✓ worked | — | — | — | `ev:llamaparse·target-2015-annual-report·complex-document-handling` | | LlamaParse | Extraction Accuracy | Bank Statement PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-bank-statement-2-jul-a511cf7bd0d6.png) | `ev:llamaparse·bank-statement-pdf·extraction-accuracy` | | LlamaParse | Extraction Accuracy | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/llamaparse-llamaparse-invoice-summary-9c05de684d8a.png) | `ev:llamaparse·invoice-pdf·extraction-accuracy` | | LlamaParse | Markdown Quality | cross-scenario | ✓ worked | — | — | — | `ev:llamaparse·cross·markdown-quality` | | LlamaParse | Markdown Quality | Financial Report | ◐ mixed | — | — | — | `ev:llamaparse·financial-report·markdown-quality` | | LlamaParse | Reading Order & Structure | Financial Report - Table Heavy | ✓ worked | — | — | — | `ev:llamaparse·financial-report-table-heavy·reading-order-structure` | | LlamaParse | Reading Order & Structure | Scanned Research Paper | ✓ worked | — | — | — | `ev:llamaparse·scanned-research-paper·reading-order-structure` | | LlamaParse | Reading Order & Structure | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:llamaparse·hybrid-earnings-report·reading-order-structure` | | LlamaParse | Reading Order & Structure | Financial Report | ✓ worked | — | — | — | `ev:llamaparse·financial-report·reading-order-structure` | | LlamaParse | Reading Order & Structure | Target 2015 Annual Report | ✓ worked | — | — | — | `ev:llamaparse·target-2015-annual-report·reading-order-structure` | | LlamaParse | Reading Order & Structure | Scanned Research Paper INT 1983-07 Issue 333 | ✓ worked | — | — | — | `ev:llamaparse·scanned-research-paper-int-1983-07-issue-333·reading-order-structure` | | LlamaParse | Schema Adherence | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/llamaparse-llamaparse-bank-statement-extracted-deta-1c069bb9dc7a.png) | `ev:llamaparse·bank-statement-pdf·schema-adherence` | | LlamaParse | Schema Adherence | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/llamaparse-llamaparse-invoice-metadata-af804b493598.png) | `ev:llamaparse·invoice-pdf·schema-adherence` | | LlamaParse | Semantic Field Enrichment | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/llamaparse-llamaparse-bank-statement-transaction-id-c01a0d30b3be.png) | `ev:llamaparse·bank-statement-pdf·semantic-field-enrichment` | | LlamaParse | Semantic Field Enrichment | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:llamaparse·invoice-pdf·semantic-field-enrichment` | | LlamaParse | Structural Clean Output | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/llamaparse-llama-extract-output-1-790fdeaa33c5.json) | `ev:llamaparse·cross·structural-clean-output` | | LlamaParse | Table & Record Completeness | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/llamaparse-llamaparse-bank-statement-transaction-it-e75317057cb2.png) | `ev:llamaparse·bank-statement-pdf·table-and-record-completeness` | | LlamaParse | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:llamaparse·invoice-pdf·table-and-record-completeness` | | LlamaParse | Table Preservation | Hybrid Earnings Report | ✓ worked | — | — | — | `ev:llamaparse·hybrid-earnings-report·table-preservation` | | LlamaParse | Table Preservation | Financial Report - Table Heavy | ✓ worked | — | — | — | `ev:llamaparse·financial-report-table-heavy·table-preservation` | | LlamaParse | Table Preservation | Scanned Research Paper | ⚠ struggled | — | — | — | `ev:llamaparse·scanned-research-paper·table-preservation` | | LlamaParse | Table Preservation | Financial Report | ◐ mixed | — | — | — | `ev:llamaparse·financial-report·table-preservation` | | LlamaParse | Table Preservation | Target 2015 Annual Report | ✓ worked | — | — | — | `ev:llamaparse·target-2015-annual-report·table-preservation` | | LlamaParse | Table Preservation | Scanned Research Paper INT 1983-07 Issue 333 | ✓ worked | — | — | — | `ev:llamaparse·scanned-research-paper-int-1983-07-issue-333·table-preservation` | | LlamaParse | Text & OCR Completeness | Scanned Research Paper | ✓ worked | — | — | — | `ev:llamaparse·scanned-research-paper·text-ocr-completeness` | | LlamaParse | Text & OCR Completeness | Scanned Research Paper INT 1983-07 Issue 333 | ✓ worked | — | — | — | `ev:llamaparse·scanned-research-paper-int-1983-07-issue-333·text-ocr-completeness` | | LlamaParse | Visual Content Retention | Hybrid Earnings Report | ✗ failed | — | — | — | `ev:llamaparse·hybrid-earnings-report·visual-content-retention` | | Nanonets | Extraction Accuracy | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-bank-statement-upi-transaction-b0f782b6c625.png) | `ev:nanonets·bank-statement-pdf·extraction-accuracy` | | Nanonets | Extraction Accuracy | Invoice PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-program-description-47fa10018717.png) | `ev:nanonets·invoice-pdf·extraction-accuracy` | | Nanonets | Schema Adherence | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-bank-statement-2-jul-a511cf7bd0d6.png) | `ev:nanonets·cross·schema-adherence` | | Nanonets | Semantic Field Enrichment | Invoice PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:nanonets·invoice-pdf·semantic-field-enrichment` | | Nanonets | Semantic Field Enrichment | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-bank-statement-2-jul-a511cf7bd0d6.png) | `ev:nanonets·bank-statement-pdf·semantic-field-enrichment` | | Nanonets | Structural Clean Output | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-bank-statement-2-jul-a511cf7bd0d6.png) | `ev:nanonets·cross·structural-clean-output` | | Nanonets | Table & Record Completeness | Bank Statement PDF | ✗ failed | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-docstrange-nanonets-transaction-count-7b5312c3e84a.png) | `ev:nanonets·bank-statement-pdf·table-and-record-completeness` | | Nanonets | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-docstrange-nanonets-extracted-line-items-5ec513388423.png) | `ev:nanonets·invoice-pdf·table-and-record-completeness` | | Reducto | Extraction Accuracy | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-invoice-summary-915cda528c4b.png) | `ev:reducto·invoice-pdf·extraction-accuracy` | | Reducto | Extraction Accuracy | Bank Statement PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-bank-statement-summary-341218e79b80.png) | `ev:reducto·bank-statement-pdf·extraction-accuracy` | | Reducto | Schema Adherence | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-bank-statement-output-320fdfd45c76.png) | `ev:reducto·bank-statement-pdf·schema-adherence` | | Reducto | Schema Adherence | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-invoice-schema-order-1938d10e3382.png) | `ev:reducto·invoice-pdf·schema-adherence` | | Reducto | Semantic Field Enrichment | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-invoice-citations-and-confidence-a61e243c4d10.png) | `ev:reducto·invoice-pdf·semantic-field-enrichment` | | Reducto | Semantic Field Enrichment | Bank Statement PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-bank-statement-schema-transactions-00d0c871c403.png) | `ev:reducto·bank-statement-pdf·semantic-field-enrichment` | | Reducto | Structural Clean Output | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-invoice-citations-and-confidence-a61e243c4d10.png) | `ev:reducto·invoice-pdf·structural-clean-output` | | Reducto | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-invoice-extracted-line0items-e23fe4fdf946.png) | `ev:reducto·invoice-pdf·table-and-record-completeness` | | Reducto | Table & Record Completeness | Bank Statement PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/reducto-reducto-transactions-5f1507b139e2.png) | `ev:reducto·bank-statement-pdf·table-and-record-completeness` | | Retab | Extraction Accuracy | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/retab-retab-invoice-summary-ab7e513edf91.png) | `ev:retab·invoice-pdf·extraction-accuracy` | | Retab | Schema Adherence | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/retab-retab-bank-statement-output-f54c1f51f37d.json) | `ev:retab·bank-statement-pdf·schema-adherence` | | Retab | Schema Adherence | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-invoice-schema-order-1938d10e3382.png) | `ev:retab·invoice-pdf·schema-adherence` | | Retab | Semantic Field Enrichment | Bank Statement PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/retab-retab-bank-statement-extracted-transacti-44a6c6e97afe.png) | `ev:retab·bank-statement-pdf·semantic-field-enrichment` | | Retab | Semantic Field Enrichment | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:retab·invoice-pdf·semantic-field-enrichment` | | Retab | Structural Clean Output | Invoice PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/datalab-invoice-schema-order-1938d10e3382.png) | `ev:retab·invoice-pdf·structural-clean-output` | | Retab | Structural Clean Output | cross-scenario | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/retab-retab-bank-statement-output-f54c1f51f37d.json) | `ev:retab·cross·structural-clean-output` | | Retab | Table & Record Completeness | Invoice PDF | ✓ worked | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/nanonets-invoice-line-item-8-ae95b671f1c5.png) | `ev:retab·invoice-pdf·table-and-record-completeness` | | Retab | Table & Record Completeness | Bank Statement PDF | ⚠ struggled | — | — | 🧾 [proof](https://d3epheqghktydj.cloudfront.net/retab-retab-bank-statement-total-transactions-c378712469f2.png) | `ev:retab·bank-statement-pdf·table-and-record-completeness` | > 🧾 = artifact-verified (proof captured) · 👁 = observed (noted, no artifact) · verdicts: worked / mixed / struggled / failed. ## The Ranking 7 tools tested head-to-head on the same input. ### 1. Landing AI — Best *Strong schema and record extraction, with weaker semantic identifier handling.* Strong schema adherence and complete row/line-item extraction, but it invents sequential transaction IDs and occasionally misreads alphanumeric codes. ### 2. Retab — Usable *Strong at nested schema reconstruction and semantic field labeling, with a few cleanup issues in record counts and final formatting.* Only tool to correctly extract transaction IDs and semantic transaction types while keeping all 51 bank-statement rows and all 8 invoice lines intact. ### 3. Extend AI — Usable *Strong structured extraction with a few cleanup gaps in ordering and field normalization.* Reliable nested JSON extraction across both documents, but transaction IDs stay blank and bank-statement totals need reconciliation. ### 4. LlamaParse — Usable *Strong at producing clean, schema-shaped JSON, but it slips on record counts and derived transaction fields.* The JSON structure is clean and the metadata is strong, but missing transaction enrichment and phantom records make it risky for finance workflows. ### 5. Datalab — Usable *Strong on schema-shaped invoice output and cited metadata, but bank-statement transaction handling breaks down on row boundaries and derived transaction fields.* Strong invoice metadata and traceability, but bank-statement rows merge together and the transaction count drifts from the source. ### 6. Reducto — Usable *Strong at reconstructing invoices and structured JSON, but the bank statement shows reliability gaps in reconciliation fields and derived transaction metadata.* The citation layer is useful, but bank-statement transaction counts inflate and the output needs post-processing before it matches the requested schema cleanly. ### 7. Nanonets — Usable *Strong schema and export handling, but weaker on bank-statement accuracy and derived-field enrichment.* Invoice metadata and export options are solid, but the bank-statement transactions are too corrupted to trust for production use. ## Full Breakdown ### Landing AI Landing AI applies the uploaded schema directly and produces clean nested JSON with a very low-friction upload, extract, and review flow. ![Landing AI screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![Landing AI screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-4-59475130ace9.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - Landing AI applied the schema cleanly, extracted all 51 bank-statement transactions, and returned complete invoice line-item and summary output with strong structural consistency. The workflow was straightforward and did not require manual tuning after schema upload. **Where it struggled:** - It replaced source transaction identifiers with sequential IDs, left some station fields null on the invoice, and misread the Ad-ID character sequence on at least one line item. **What came out:** ![Landing AI output showing Bank statement metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-extracted-bank-statement-00acbc029c1e.png) *Output — The JSON extraction shows statement metadata and account-holder address details, including Standard Chartered, the 16 Jul 2019 statement date, INR currency, and the full Chennai address.* ![Landing AI output showing Bank statement account details](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-extracted-bank-statement-acco-86965ed85143.png) *Output — The account, branch, statement period, and opening/closing balances are extracted into separate structured objects with the expected dates and amounts.* ![Landing AI output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-transaction-items-602d1d5357b9.png) *Output — The transactions array is expanded and shows that Landing AI preserved the full transaction list as structured JSON records.* ![Landing AI output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-bank-statement-18-jun-extract-c6073a9ae222.png) *Output — The 18 Jun transaction extraction maps each row into transaction_id, date, value_date, description, transaction_type, amount, and running balance fields.* ![Landing AI output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-bank-statement-18-jun-extract-0c54dcb64c1b.png) *Output — The annotated view highlights transaction_id fields, emphasizing that the tool assigns extracted IDs to every transaction record.* ![Landing AI output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-extracted-invoice-metadata-57a291c39e3d.png) *Output — The invoice metadata output contains invoice_number, dates, advertiser, and product fields in structured JSON.* ![Landing AI output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-invoice-extracted-line-items-1e5e7030716e.png) *Output — The invoice output shows the line_items array expanded with entries 0 through 7, confirming that all eight ad records are present.* ![Landing AI output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-invoice-summary-c320d8e6d1e1.png) *Output — The summary returns aired_spots 8, gross_total 29750, agency_commission 4462.5, net_amount_due 25287.5, and payment_terms 30 Days.* ![Landing AI output showing Invoice station fields](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-invoice-station-e066b6555b5f.png) *Output — The station object keeps call_letters but leaves address and several contact fields null, even though the same information appears elsewhere in the invoice output.* ![Landing AI output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-extracted-line-item-1-6f5a3b324a8b.png) *Output — The first invoice line item is structured correctly and includes the Ad-ID field populated with the extracted code.* ![Landing AI output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-landing-ai-invoice-ad-id-line-item-1-3637c88c5d00.png) *Output — The highlighted Ad-ID value NRCCW1071005 matches the source row, but the code itself contains an OCR error because the source value should read NRCCWI071005.* ### Retab Retab uses a node-based PDF input plus Extract workflow where the schema is pasted directly into the extractor, then returns copyable structured JSON. ![Retab screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![Retab screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-3-f56b3214f4cf.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - Retab reconstructed the bank statement schema cleanly, extracted all 51 transaction rows, derived transaction_id values from embedded description strings, and classified transaction types accurately. On the invoice, it extracted all 8 line items and preserved the financial totals without record loss. **Where it struggled:** - The bank-statement summary.total_transactions value was 43 instead of the expected 40. On the invoice, payment_terms retained label text and the JSON key order differed from the supplied schema. **What came out:** ![Retab output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-bank-statement-extracted-transacti-bb4bdef8ec56.png) *Output — The extracted transaction record includes a real transaction_id value and classifies the entry as UPI, showing that Retab derives IDs from the description text instead of leaving them blank.* ![Retab output showing Bank statement summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-bank-statement-total-transactions-d4db7cb55e38.png) *Output — The summary reports total_transactions as 43, which is higher than the expected 40 after excluding balance-forward, tax, and charge entries.* ![Retab output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-extracted-invoice-metadata-1edee57417fc.png) *Output — The invoice metadata output includes the expected advertiser, billing, account details, flight dates, and invoice identifiers in structured JSON.* ![Retab output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-invoice-extracted-line-items-d98bbb0dffcc.png) *Output — The line_items array is expanded and shows the extracted invoice line records as separate structured entries.* ![Retab output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-invoice-extracted-line-item-8eaca32c81a7.png) *Output — The extracted line item preserves channel, description, air date, air time, rate, and flight-period fields for a single ad placement.* ![Retab output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-invoice-summary-ccf2d87a6aeb.png) *Output — The invoice summary returns the expected aired_spots, gross_total, agency_commission, net_amount_due, and payment_terms values.* ![Retab output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-invoice-summary-2-413f904fb43a.png) *Output — The summary repeats the same totals, but payment_terms keeps the label text and reads as Payment Terms 30 Days instead of just 30 Days.* ![Retab output showing Schema order](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-retab-invoice-schema-order-634c24eba321.png) *Output — The extracted JSON remains valid, but the key order differs from the order defined in the supplied invoice schema.* ### Extend AI Extend AI returns structured JSON with confidence metadata and good nested coverage, but some field-level cleanup is still needed before production use. ![Extend AI screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![Extend AI screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-8-f348330571d1.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - Extend AI populated the bank-statement metadata, balances, rewards, and transaction rows cleanly, and it also returned invoice line items and totals in a nested JSON structure. The confidence metadata is useful for review workflows. **Where it struggled:** - It left transaction_id blank on the bank statement, reported a mismatched transaction total, reordered some schema keys, and kept label text inside payment_terms. **What came out:** ![Extend AI output showing Bank statement metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-bank-statement-extracted-detai-b6e6aca0df93.png) *Output — The bank statement JSON includes branch, account, and rewards structures, showing that the tool fills nested schema objects rather than returning flat OCR text.* ![Extend AI output showing Bank statement document view](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-document-level-bank-statement-6a3ae27f2afa.png) *Output — The document-level output includes summary totals and opening/closing balance dates alongside metadata and rewards fields.* ![Extend AI output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-extracted-transaction-item-3536dd55477d.png) *Output — The extracted transaction record preserves date, value_date, description, amount, running balance, and transaction_type for an individual bank-statement row.* ![Extend AI output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-transaction-type-b55e6eb13b26.png) *Output — The annotated transaction snippet shows the tool classifying the record as Withdrawal based on the description text.* ![Extend AI output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-extracted-transaction-id-a877374aa6b1.png) *Output — The transaction_id field remains null even though the description contains an embedded reference number.* ![Extend AI output showing Bank statement summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-bank-statement-summary-77b88864e2a2.png) *Output — The summary reports total_transactions as 49, which does not reconcile with the expected count after excluding balance-forward, tax, and charge entries.* ![Extend AI output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extendai-extracted-invoice-line-items-fdffc2d2b904.png) *Output — The invoice line_items array is expanded and shows the eight structured advertising records.* ![Extend AI output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-invoice-summary-1e3fd33189a1.png) *Output — The invoice summary contains the expected aired_spots, gross_total, net_amount_due, and agency_commission values, but payment_terms still includes the label text.* ![Extend AI output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-invoice-summary-2-ef8a0e2dbc9c.png) *Output — The summary output again shows the payment terms as Payment Terms 30 Days instead of only the value 30 Days.* ![Extend AI output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-invoice-line-item-6-b5ded30b0feb.png) *Output — A single invoice line item is extracted with rate, ad_id, time slot, flight dates, and day pattern preserved in JSON.* ![Extend AI output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-invoice-extracted-metadata-eb058eb465d2.png) *Output — The metadata view shows confidence values for station, summary, advertiser, and line_items sections, which helps validation before export.* ![Extend AI output showing Schema order](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-extend-ai-schema-order-invoice-b160c84f84be.png) *Output — The invoice output is structurally valid but the object order differs from the schema order defined by the user.* ### LlamaParse LlamaParse reconstructs the schema cleanly and exposes confidence information, but it misses key transaction enrichment fields and introduces phantom records. ![LlamaParse screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![LlamaParse screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-7-0f2e00888086.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - LlamaParse reconstructs the nested schema well and keeps the output cleanly structured, with metadata and line items exposed in a reviewable JSON format. The invoice summary values are also accurate. **Where it struggled:** - Transaction IDs and transaction types are empty, several value_date fields are missing, the bank statement transaction count is too high, and the invoice output introduces a phantom ninth line item. **What came out:** ![LlamaParse output showing Bank statement metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-bank-statement-extracted-deta-31e50a30ba6b.png) *Output — The bank-statement details panel reconstructs metadata, account holder, account, branch, statement period, balances, transactions, summary, rewards, and disclaimers.* ![LlamaParse output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-bank-statement-transaction-it-526f12ceb96d.png) *Output — The transaction record preserves the amount and balance for a UPI-style deposit, but the transaction_id and transaction_type fields are empty.* ![LlamaParse output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-bank-statement-transaction-id-944dde08fbc1.png) *Output — The highlighted transaction snippet shows transaction_id and transaction_type both left blank on an ATM withdrawal record.* ![LlamaParse output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-bank-statement-16-jul-value-d-4906fee56b7b.png) *Output — The highlighted statement row shows a value-date style layout, but the extracted transaction records still leave value_date empty in some cases.* ![LlamaParse output showing Bank statement summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-bank-statement-summary-5ddaf0f99859.png) *Output — The summary reports total_deposits, total_withdrawals, and total_transactions, but the transaction total does not align with the source document.* ![LlamaParse output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-bank-statement-transaction-co-e438230e98a1.png) *Output — The tree view shows the transactions array extending into the high 40s and 50s, consistent with the over-counted transaction list.* ![LlamaParse output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-invoice-metadata-d0f19e292f47.png) *Output — The invoice metadata output includes invoice, advertiser, and station fields such as invoice number, date, period, and station contact details.* ![LlamaParse output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-invoice-extracted-line-items-e5a8eb8bdd8e.png) *Output — The invoice line_items array is expanded in the viewer and shows entries 0 through 7.* ![LlamaParse output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/document-extraction-llamaparse-invoice-extracted-line-item-2-3dc7cd242bee9383.png) *Output — The line item record preserves the 9:38 AM airing, ad_id, rate, time slot, and flight period for line number 2.* ![LlamaParse output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-llamaparse-invoice-summary-b84bf714321e.png) *Output — The invoice summary returns the expected aired_spots, gross_total, agency_commission, net_amount_due, and payment_terms values.* ![LlamaParse output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/document-extraction-llamaparse-invoice-extracted-line-item-8-dcb6d19316086e7e.png) *Output — The output shows a phantom line number 9 beginning below line number 7, even though the source invoice contains only 8 advertising records.* ### Datalab Datalab adds citations for each field and is especially strong on invoice extraction, but the bank statement shows segmentation and count issues. ![Datalab screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![Datalab screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-5-2cd43fae8cf4.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - Datalab produced reliable invoice metadata and line-item extraction with citation traceability, and the invoice totals matched the source document. The field-level citations are useful for audit and validation workflows. **Where it struggled:** - On the bank statement, several rows were merged, transaction boundaries shifted, transaction classification was inconsistent, and the count ended up at 54 instead of the expected 51. **What came out:** ![Datalab output showing Bank statement metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-bank-statement-extracted-details-01f74ce0e390.png) *Output — The bank-statement metadata output shows the bank name, statement date, currency, and account-holder name with citation paths attached.* ![Datalab output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-bank-statement-transaction-extra-264c546c8d32.png) *Output — The transaction record includes date, value date, description, transaction type, withdrawal amount, and running balance with citations.* ![Datalab output showing Bank statement citation output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-bank-statement-extracted-citatio-070e96d8c0fd.png) *Output — The extracted disclaimer field includes source citations for the insurance coverage and reporting period text.* ![Datalab output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-bank-statement-21-jun-extracted--d94e7d9b89ad.png) *Output — The 21 Jun extraction shows a long merged UPI and merchant description, which indicates that transaction boundaries were not preserved cleanly.* ![Datalab output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-bank-statement-extracted-28-jun--04644cff3353.png) *Output — The 28 Jun row is extracted as a Deposit even though the row contains a withdrawal amount, showing a classification error.* ![Datalab output showing Bank statement summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-bank-statement-transaction-count-dc5e520a0f17.png) *Output — The summary reports total_transactions as 54 and includes citation paths, but the count does not match the source statement.* ![Datalab output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-invoice-metadat-d7a0d73aa4b5.png) *Output — The invoice metadata block includes invoice identifiers, period dates, order number, and estimate number with citations from the source pages.* ![Datalab output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-invoice-line-items-91919439c1ab.png) *Output — The line_items array is expanded and shows all eight advertising line items in the invoice output.* ![Datalab output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-invoice-line-item-2-ab8a8332ed35.png) *Output — The line item record includes channel, description, time slot, day of week, air date, length, air time, and ad ID with citations.* ![Datalab output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-invoice-summary-7616e51aac27.png) *Output — The invoice summary returns aired_spots, gross_total, agency_commission, net_amount_due, and payment_terms with citation references.* ![Datalab output showing Schema order](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-datalab-invoice-schema-order-1c5896492113.png) *Output — The output is structurally valid, but the schema order is not preserved in the generated JSON.* ### Reducto Reducto adds citations, bounding boxes, and granular confidence, but the bank-statement counts drift and the output needs transformation before it matches a clean schema. ![Reducto screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![Reducto screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-6-689a2395ada2.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - Reducto gives a detailed audit trail with citations, bounding boxes, and confidence metadata, and it preserves the invoice's repeated line items and totals accurately. The nested structure is recognizable and reviewable. **Where it struggled:** - The bank statement returns too many transaction records, summary totals are wrong, and several transaction-level fields are missing. The output also carries extra metadata that is not part of the requested schema, so it needs a transformation step before direct consumption. **What came out:** ![Reducto output showing Bank statement metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-bank-statement-output-6745bf8903b7.png) *Output — The parsed bank statement output reconstructs metadata, account-holder details, branch information, and the statement period in nested JSON.* ![Reducto output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-transactions-61770cc60bde.png) *Output — The transactions tree shows entries in the mid-40s through low-50s, but the count is inflated relative to the source statement.* ![Reducto output showing Bank statement summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-bank-statement-summary-f62d77d1a0a4.png) *Output — The summary reports total_transactions as 70, which is far above the expected total for the statement.* ![Reducto output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-extracted-transactions-af4753c92b3a.png) *Output — The extracted transaction array shows dates, value dates, descriptions, and balances, but several transaction-level fields are missing or incomplete.* ![Reducto output showing Bank statement confidence output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-bank-statement-confidence-scores-1f132796379b.png) *Output — The confidence view shows a transaction field with high confidence and granular parse confidence metadata.* ![Reducto output showing Bank statement citation output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-bank-statement-citation-annotate-b15869a7cc4e.png) *Output — The annotated extraction highlights the citation object tied to the extracted customer name, showing the audit-trail style output.* ![Reducto output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-invoice-metadata-683485bb60c5.png) *Output — The invoice metadata output includes a detected invoice number with citation details and bounding-box metadata.* ![Reducto output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-invoice-extracted-line0items-bcab7a05d9b9.png) *Output — The invoice line_items tree shows entries 0 through 7 and reconstructs the repeating ad rows in structured JSON.* ![Reducto output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-invoice-summary-069e92d5007b.png) *Output — The invoice summary returns the expected aired_spots, gross_total, agency_commission, net_amount_due, and payment_terms values.* ![Reducto output showing Invoice confidence output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-invoice-citations-and-confidence-a61e243c4d10.png) *Output — The field-level confidence view shows citations and granular parse confidence for an invoice field, demonstrating the audit layer.* ![Reducto output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-reducto-invoice-additional-field-bcab4298fc52.png) *Output — The output also includes extra fields such as idb_number and citation metadata, making the JSON more verbose than the requested schema.* ### Nanonets Nanonets has strong document-level metadata extraction and broad export options, but its bank-statement transaction parsing is too corrupted to be reliable. ![Nanonets screenshot showing Bank statement input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-pdf-2-f607ef6b9fb6.pdf) *Screenshot — The bank statement PDF used for the structured extraction test.* ![Nanonets screenshot showing Invoice input](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-pdf-2-b80234078612.pdf) *Screenshot — The invoice PDF used for the structured extraction test.* **What worked:** - Nanonets extracted the invoice metadata cleanly, returned all eight invoice line items as separate records, and supported multiple export formats. The invoice summary numbers also reconcile correctly. **Where it struggled:** - The bank statement is not reliable: transaction descriptions are heavily garbled, several dates are missing, transaction_type is null across the board, and transaction_id is null. The invoice descriptions also merge labels into the field text, and the eighth line item is partially incomplete. **What came out:** ![Nanonets output showing Bank statement metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-bank-statement-983e48e20353.png) *Output — The bank statement output follows the schema and reconstructs the higher-level nested sections instead of returning flat OCR text.* ![Nanonets output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-bank-statement-trans-93295dbbfc56.png) *Output — The transaction viewer shows separate entries in the array, but the descriptions are already starting to look merged and noisy.* ![Nanonets output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-extracted-bank-state-353f2eff4036.png) *Output — The transaction row content is badly concatenated, with multiple descriptions merged into one record and transaction_type left null.* ![Nanonets output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-bank-statement-18-jun-transactions-7b31e30bb54d.png) *Output — The 18 Jun source rows show distinct ATM withdrawal, purchase, and deposit entries, which the extraction should have preserved separately.* ![Nanonets output showing Bank statement transaction output](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-bank-statement-dates-7d868a5e9753.png) *Output — The bank-statement date view shows that dates are missing on many extracted transactions even though the source rows contain them.* ![Nanonets output showing Bank statement summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-extracted-total-tran-166ed6f76bd4.png) *Output — The summary reports total_deposits and total_withdrawals but total_transactions is null.* ![Nanonets output showing Invoice metadata](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-extracted-metadata-70078dca1703.png) *Output — The invoice metadata output correctly extracts the invoice, advertiser, station, account, billing, and remit sections.* ![Nanonets output showing Invoice line items](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-extracted-line-items-3f010d95c9da.png) *Output — The line_items tree shows all eight ad records present in the invoice output.* ![Nanonets output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-line-item-4-236f72aac69e.png) *Output — Line item 4 is extracted with channel, description, time slot, air date, ad ID, rate, and schedule fields.* ![Nanonets output showing Invoice summary](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-invoice-summary-6927b6057172.png) *Output — The invoice summary returns 8 aired spots, gross_total 29750, agency commission 4462.5, net amount due 25287.5, and payment terms of 30 days.* ![Nanonets output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-invoice-program-description-b2e586bed996.png) *Output — The source invoice row shows the program description and the adjacent political issue label, which should have been separated cleanly in the extraction.* ![Nanonets output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-extracted-program-de-dbf7cfcf3a71.png) *Output — The extracted description merges the program title and rate label into one string, reducing description cleanliness.* ![Nanonets output showing Invoice line item](https://d3epheqghktydj.cloudfront.net/extract-and-query-structured-data-from-d-docstrange-nanonets-line-item-8-0eb9b4503cb7.png) *Output — Line item 8 is only partially extracted, with flight-period and frequency fields missing even though the source row contains a complete record.* ## Final Take Landing AI is the overall winner here, and the scorecards support that: it combines perfect schema adherence, extraction accuracy, table-and-record completeness, and clean output, making it the strongest all-around choice for schema-faithful record extraction. The main caveat is that its semantic field enrichment is weaker than some rivals, so it is less compelling when identifier interpretation or field labeling is the priority. If you need the best cleanup and completeness balance, Landing AI is the safest pick; if you need stronger semantic labeling, Retab is the more specialized option, though it gives up record completeness and final formatting. LlamaParse and Extend AI are also credible for structured output, but both show more trade-offs than Landing AI: LlamaParse is very clean but slips on record counts and derived transaction fields, while Extend AI has strong structure but lower extraction accuracy and some normalization/ordering cleanup gaps. Datalab, Nanonets, and Reducto trail the top group because their transaction handling, enrichment, or accuracy is less consistent in the scorecards. Tested as of June 2026 · re-verified monthly.