Filter records across documents
Filter a subset of the tool's extracted records across multiple documents, using one extracted record per source document as the operating grain.
What this scenario means
This scenario checks whether the tool can work over the records it has already extracted, rather than rereading source documents one by one. A good system applies a condition across the extracted set, keeps the document-level grain intact, and returns only the matching records. The hard part is that a tool can expose records without truly supporting cross-document querying, so this test separates record extraction from record selection.
What we evaluate
- Whether the tool filters records from its own extracted set across multiple documents.
- Whether the result keeps one extracted record per source document as the operating grain.
- Whether only matching records are returned and non-matching records are excluded.
- Whether the operation happens inside the tool over extracted records rather than by exporting data elsewhere.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Querying
Operates over the extracted records — filters, totals, answers questions — and is clear when an answer didn't come from them.
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
Structured Document Extraction
Which document extraction platform turns a user-defined schema into correct structured records — across layouts, scans and repeated sets — and is honest about what it could not find?