The output#1
docling
It produced a clean result for the financial report overall, with strong text, hierarchy, and most tables, but it still dropped the visual assets and flattened one complex table header.
docling-docling-input2-financialpdf-fy2025-heade-5340c1f4d35a.png
The output#2
liteparse
Clean prose, headings, and table-of-contents handling come through well, but several tables flatten or split apart, which keeps it from being a top score.
liteparse-liteparse-input2-financialpdf-summary-se-75cdedba0497.png
The output#3
pymupdf4llm
It can pull through some narrative sections, but the report’s core tables disappear, the header/logo becomes garbled text, and reading order and overall structure become unreliable.
pymupdf4llm-pymupdf4llm-input2-financialpdf-quarterl-052a903bfa21.png
The output#4MDdoc2mark-doc2mark-input2-financialpdf-output-64777dd9b209.mdopen raw ↗ doc2mark
It gets the plain text and contents order right, but table formatting and heading consistency are shaky.
doc2mark-doc2mark-input2-financialpdf-output-64777dd9b209.md
The output#5MDmarkitdown-markitdown-input2-financialpdf-output-1b08db40e794.mdopen raw ↗ markitdown
It is solid for plain text and the table of contents stays in order, but the financial tables and section structure are too uneven to call it reliable for structured extraction.
markitdown-markitdown-input2-financialpdf-output-1b08db40e794.md
Comments (0)