Benchmark · Version 1

Structured Document Extraction

Evaluates whether a document extraction platform can turn a user-defined schema into correct structured records across layouts, scans and repeated sets, while staying honest about what it could not find.

What this benchmark is

This benchmark evaluates document extraction platforms that take a user-defined schema and return structured records. It asks whether a product can handle ordinary documents, scans and mixed pages, repeated sets, training from labelled examples, corrections and record queries.

A buyer learns how far the platform can go in turning documents into usable structured output, and whether it clearly separates extracted facts from missing ones or values found elsewhere in the document.

This benchmark uses the registered round-1 PDF fixture corpus.

In scope

  • User-defined schema to structured records from many documents
  • Scalar fields and repeated or nested records
  • Digital, scanned and mixed documents
  • Training from labelled examples
  • Correcting returned data
  • Working with extracted records: filtering, totalling and asking questions

Out of scope

  • Document classification
  • Packet splitting
  • Routing document types to different schemas
  • Transcription quality
  • Non-English documents
7Capabilities
19Scenarios

Capabilities included

The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.

CapabilityWhat it means hereScenarios
Field ExtractionReturns the scalar fields the schema asks for — present, interpreted, derived, absent or differently laid out — correctly typed5
Nested & Repeated FieldsReturns repeated sets of records completely, with each value under the right record: none dropped, none invented, none cross-bound4
TrainingLearns from labeled examples — documents paired with their expected output — and applies what it learned to a document it hasn't seen1
OCRThe same schema returns the same values when the document is an image, and a mixed document is handled page by page2
Source GroundingPoints to where a returned value came from — and attaches no pointer to a value that wasn't there2
Review and CorrectionA human fix reaches the stored record and the systems downstream of it3
QueryingOperates over the extracted records — filters, totals, answers questions — and is clear when an answer didn't come from them3

Scenarios included

A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Scoring

Unsupported capabilities

A capability a tool doesn't support scores 0, and the 0 stays in the denominator. No tool ranks higher for having less product.

Methodology, weighting and ranking were still open in the v1 design pass.

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.