--- title: "Dolphin (ByteDance)" type: "AI Tool" url: "https://aidemos.com/tools/dolphin" description: "On a scanned paper, Dolphin-1.5 preserved reading order, captions, text, and routine tables, but dropped charts and misread repeated images." category: "developer-tools" website: "https://github.com/bytedance/Dolphin" published: "2026-08-19T15:54:38.579133+00:00" updated: "2026-08-19T16:05:31.913556+00:00" --- # Dolphin (ByteDance) Self-hosted PDF-to-Markdown that preserves most text and tables, but drops charts and struggles with repeated images. ## TL;DR Verdict **Strong baseline for text and tables, but not for figures** **Where it wins:** - You want a free, self-hosted PDF-to-Markdown baseline that usually preserves prose, headings, tables, captions, and reading order. - Your PDFs are mostly text-heavy or table-heavy, and you can tolerate some manual cleanup for charts and repeated images. - You care more about data values and document structure than about faithfully recovering chart graphics. **Main limitation:** You need chart data or figure content preserved in markdown; every chart tested was reduced to a placeholder. **Pricing:** Open source Free `Open source` · `PDF → Markdown` · `Scanned OCR` · `Charts omitted` **Website:** [Visit Dolphin (ByteDance)](https://github.com/bytedance/Dolphin) > **Strong baseline for text and tables, but not for figures** > > Dolphin-1.5 is a solid lightweight self-hosted baseline for PDF-to-Markdown when the job is mostly prose, headings, and routine tables. It handled the scanned paper well at the OCR level and kept reading order/captions intact, but it consistently dropped chart content and behaved inconsistently around repeated images, signatures, and some complex table structures. ## Demo Recording [Video: Dolphin (ByteDance) demo recording](https://cdn.futuresmart.ai/public/aidemos/40372936300d450383b4b362312af98a.mp4?v=1) *Video — Colab notebook walkthrough showing Dolphin-1.5 loading the model, uploading a PDF, and previewing the extracted markdown.* ## Feature-by-Feature Breakdown ### Embedded Visual Element Retention **Verdict:** Omitted Dolphin preserves embedded non-text elements from chart-heavy pages in hybrid earnings reports and scanned research papers, plus report-signature pages with images, logos, and handwritten signatures, often leaving them as Figure placeholders while keeping nearby captions and annotations in sequence. The sampled cases show that chart content itself is not reconstructed from data. **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** **Output:** **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Bottom line:** Not safe for chart data extraction; nearby captions and annotations can survive, but the chart itself does not. ### OCR and Text Extraction **Verdict:** Mostly accurate, with recurring glyph errors Dolphin extracts body text from native and scanned PDFs, including a two-column title page, but the sampled cases include misreads on special characters, checkbox glyphs, dash substitution, and repeated header text. **Input:** **Output:** **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** **Output:** **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** **Output:** **Bottom line:** Good prose OCR overall, but special characters and some repeated header text need review. ### Table Reconstruction **Verdict:** Mixed Dolphin reconstructs tables from financial and research PDF layouts, including balance-sheet and business-results tables with merged headers and page-spanning splits, while the harder samples expose problems with wrapped labels and deeply nested row groups. **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** **Output:** **Bottom line:** Strong on routine tables; fragile on wrapped labels and deeply nested layouts. ### Reading Order Preservation **Verdict:** Preserved Dolphin preserves logical content sequence in sampled two-column pages and page-boundary table splits, keeping nearby paragraphs and sections in the intended order. **Input:** > **Image** **Output:** **Input:** **Output:** **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** Source crop > **Image** — Source crop **Output:** Rendered output > **Image** — Rendered output **Input:** > **Image** **Output:** **Bottom line:** Sequencing held up in all sampled cases. ### Document Hierarchy Detection **Verdict:** Inconsistent Dolphin emits markdown heading structure for section titles in the hybrid report, but the sampled headings are not always assigned consistent levels or any level at all. **Input:** **Output:** **Input:** **Output:** **Input:** **Output:** **Input:** **Output:** **Input:** **Output:** **Bottom line:** Usable, but heading levels are not reliable in the mixed-layout filing. ### Caption and Footnote Association **Verdict:** Preserved Dolphin keeps captions, titles, dates, and table notes attached to the correct figure or table in the sampled documents, even when the parent image or table is degraded. **Input:** **Output:** **Input:** **Output:** **Input:** **Output:** **Bottom line:** Association stayed intact in the checked samples. ## Free to self-host MIT-licensed Dolphin-1.5; the report says it is free to self-host. | Plan | Price | Notes | | --- | --- | --- | | Open source ★ | Free | MIT license; self-hosted deployment. | *The evaluation also notes a newer Dolphin-v2 exists, but it was not the version tested here.* ## Is It Right For You? **Use it if** - You want a free, self-hosted PDF-to-Markdown baseline that usually preserves prose, headings, tables, captions, and reading order. - Your PDFs are mostly text-heavy or table-heavy, and you can tolerate some manual cleanup for charts and repeated images. - You care more about data values and document structure than about faithfully recovering chart graphics. **Skip it if** - You need chart data or figure content preserved in markdown; every chart tested was reduced to a placeholder. - Your documents rely on repeated images, handwritten signatures, or exact checkbox glyphs that must be consistent. - You need highly complex wrapped-label or nested tables to stay perfectly aligned without manual review. ## Classification - **Category:** developer-tools - **Subcategory:** documentation-tools - **Type:** text - **Built for:** Other ## Frequently Asked Questions **Q: Does Dolphin preserve charts in PDFs?** Not in the tested documents. All charts in the hybrid earnings report and the scanned research paper were reduced to Figure placeholders, with no chart data recovered. **Q: How well does Dolphin handle tables?** It handled many routine financial tables well, including merged headers and a page-spanning balance sheet split. But wrapped labels, diagonal-split headers, and the most complex nested table caused structural errors. **Q: Can Dolphin OCR scanned PDFs?** Yes, at least at the prose level. The scanned research paper was OCR'd well overall, including a two-column title page, though charts were omitted and the most complex table still broke down. **Q: Does Dolphin preserve reading order in two-column layouts?** Yes in the sampled case. The split paragraph on the hybrid earnings report and the two-column title page in the scanned paper were linearized into sensible reading order. **Q: Is Dolphin free to use?** Yes. The report says Dolphin-1.5 is open source under the MIT license and free to self-host. ## Similar Tools AI tools similar to Dolphin (ByteDance): - [docling](https://aidemos.com/tools/docling) — Open-source PDF-to-markdown conversion that is strong on text, headings, and standard tables, but drops charts and other visual assets. - [PyMuPDF4LLM](https://aidemos.com/tools/pymupdf4llm) — Open-source PDF-to-markdown for clean native PDFs, but unreliable on scans, dense tables, and images. - [liteparse](https://aidemos.com/tools/liteparse) — Open-source PDF-to-markdown parsing that works well on native-digital reports, but degrades on scans, tables, and charts. - [doc2mark](https://aidemos.com/tools/doc2mark) — Open-source PDF-to-markdown that preserves native text and headings, but still struggles with tables, charts, images, and scans. - [markitdown](https://aidemos.com/tools/markitdown) — Fast native-PDF text extraction for markdown, but structure, charts, images, and scans are unreliable. - [MinerU](https://aidemos.com/tools/mineru) — Open-source PDF-to-markdown that keeps figures and charts in place, but tables and punctuation can get messy on complex files. - [PaddleOCR](https://aidemos.com/tools/paddleocr) — Open-source PDF-to-markdown conversion that keeps images and tables in place, but still corrupts text and complex structure on harder documents. ## Need a custom AI solution for this use case? If you are looking to build a custom PDF-to-Markdown conversion, document parsing, or structured extraction system for your business or internal workflow, email us at [contact@futuresmart.ai](mailto:contact@futuresmart.ai). ### Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at [collaborate@aidemos.com](mailto:collaborate@aidemos.com).