Structured Document Extraction
Evaluates whether a document extraction platform can turn a user-defined schema into correct structured records across layouts, scans and repeated sets, while staying honest about what it could not find.
What this benchmark is
This benchmark evaluates document extraction platforms that take a user-defined schema and return structured records. It asks whether a product can handle ordinary documents, scans and mixed pages, repeated sets, training from labelled examples, corrections and record queries.
A buyer learns how far the platform can go in turning documents into usable structured output, and whether it clearly separates extracted facts from missing ones or values found elsewhere in the document.
This benchmark uses the registered round-1 PDF fixture corpus.
In scope
- User-defined schema to structured records from many documents
- Scalar fields and repeated or nested records
- Digital, scanned and mixed documents
- Training from labelled examples
- Correcting returned data
- Working with extracted records: filtering, totalling and asking questions
Out of scope
- Document classification
- Packet splitting
- Routing document types to different schemas
- Transcription quality
- Non-English documents
Capabilities included
The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.
| Capability | What it means here | Scenarios |
|---|---|---|
| Field Extraction | Returns the scalar fields the schema asks for — present, interpreted, derived, absent or differently laid out — correctly typed | 5 |
| Nested & Repeated Fields | Returns repeated sets of records completely, with each value under the right record: none dropped, none invented, none cross-bound | 4 |
| Training | Learns from labeled examples — documents paired with their expected output — and applies what it learned to a document it hasn't seen | 1 |
| OCR | The same schema returns the same values when the document is an image, and a mixed document is handled page by page | 2 |
| Source Grounding | Points to where a returned value came from — and attaches no pointer to a value that wasn't there | 2 |
| Review and Correction | A human fix reaches the stored record and the systems downstream of it | 3 |
| Querying | Operates over the extracted records — filters, totals, answers questions — and is clear when an answer didn't come from them | 3 |
Scenarios included
A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.
Field Extraction
5 scenariosNested & Repeated Fields
4 scenariosSource Grounding
2 scenariosHow the benchmark works
A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.
Unsupported capabilities
A capability a tool doesn't support scores 0, and the 0 stays in the denominator. No tool ranks higher for having less product.
Methodology, weighting and ranking were still open in the v1 design pass.
Resources and fixtures
The registered material and systems that create a consistent test environment for this benchmark.
Structured Document Extraction fixture corpus — round 1
A fixed round-1 fixture corpus of 19 fictional PDFs for Structured Document Extraction (R5 v1).