The set continues across a page break
A repeated record set continues across a page boundary and should stay one array.
What this scenario means
This tests whether an extraction tool preserves a repeated set when the records spill past a page boundary. The hard part is keeping the set continuous and complete, rather than splitting it into separate arrays or introducing boundary duplicates.
What we evaluate
- Whether the tool returns the whole repeated set as one array.
- Whether it avoids splitting the set into separate arrays at the boundary.
- Whether it avoids dropping members that appear after the boundary.
- Whether it avoids duplicating members at the boundary.
- Whether each member remains associated with the correct record across the full set.
Capabilities this scenario exercises
A scenario may exercise one or more capabilities.
Nested & Repeated Fields
Returns repeated sets of records completely, with each value under the right record: none dropped, none invented, none cross-bound. ON TRIAL: merges back into Field Extraction if precision and recall do not actually diverge on repeated sets.
Benchmarks that use this scenario
A scenario has global identity and may be reused across benchmarks.
Structured Document Extraction
Which document extraction platform turns a user-defined schema into correct structured records — across layouts, scans and repeated sets — and is honest about what it could not find?