Benchmark · Version 1

Converting a complex PDF into clean Markdown with an open-source library

How well a self-hostable open-source library turns complex PDFs into faithful Markdown that a reader can trust without opening the original.

What this benchmark is

This benchmark is for self-hostable open-source libraries that convert real-world PDFs into Markdown. It helps a buyer judge whether a tool keeps the source document usable when it contains native text, tables, figures, charts, scanned pages, equations, or code.

A reader learns where a library preserves content and structure, and where it drops, reorders, invents, or weakens information. The point is not pretty output; it is whether the converted document still supports search, retrieval, QA, indexing, and human review.

Each test uses the complete PDF, not a cropped page. The tool runs once with its default pipeline and no per-tool flag tuning, and the expected source truth is recorded from the PDF before the run. The grader checks the relevant page or region in the output, not the whole document at once.

In scope

  • Converting complex, real-world PDFs into Markdown with an open-source library.
  • Preserving text, reading order, headings, tables, figures, charts, scanned text, equations, and code.
  • Judging completeness, accuracy, faithfulness, and honesty against the source PDF.

Out of scope

  • Structured field extraction.
  • Web-page conversion.
  • Form filling.
  • Handwriting.
  • Document classification or routing.
  • PDF creation or editing.
8Capabilities
12Scenarios

Capabilities included

The broad abilities this benchmark evaluates. Each capability is defined globally; this page states that it belongs to this benchmark.

CapabilityWhat it means hereScenarios
Text FidelityWhether ordinary digital text comes through complete, unchanged, and not invented in the Markdown output.1
Reading Order & LayoutWhether the source page is read in the natural sequence, so text is not interleaved or reordered by layout.1
Heading & Section StructureWhether headings and subheadings become real Markdown structure, with section hierarchy intact and footnotes kept attached outside the body flow.2
Table ExtractionWhether tables keep their rows, columns, headers, and cell values as a table.2
Figures & ChartsWhether figures and charts remain placed, referenced, and captioned or titled, with a text trace that still points back to the source.2
Scanned Document OCRWhether image-only pages are turned into accurate text instead of being skipped or left unread.2
Equations & Mathematical NotationWhether equations survive as math markup or a clear fallback instead of garbled prose.1
Code ExtractionWhether code survives as a fenced block with line breaks and indentation intact.1

Scenarios included

A scenario is a real-world situation used to test a capability. Together these scenarios define the evaluation scope of version 1.

How the benchmark works

A public summary of the evaluation method. The same defined scope and evidence standard apply to every tool assessed under this version.

Evaluation unit

One test case for one tool

Each verdict applies to one test case run against one tool.

Outcomes

Pass, fail, or not gradable

A result is pass, fail, or not gradable; not gradable is not a failure.

Consistency

Versioned scenarios and registered resources

Use the registered scenario versions and the registered resource versions for every run.

Evidence standard

Recorded stimulus, complete output, relevant references

Keep the recorded stimulus, the complete output, and the relevant registered references with the verdict.

Active test cases are partly withheld to reduce benchmark gaming; capabilities, scenarios, method, and resource types stay public.

Resources and fixtures

The registered material and systems that create a consistent test environment for this benchmark.

Converting a complex PDF into clean Markdown with an open-source library — Benchmark definition | AI Demos