PulseAugur
实时 16:09:27
English(EN) I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types

MinerU、Granite-Docling、PaddleOCR-VL 的 PDF 解析能力对比

一位用户对 MinerU、Granite-Docling 和 PaddleOCR-VL 这三个 PDF 解析模型进行了比较分析,评估了它们在 12 项不同能力和 6 种不同文档类型上的表现。测试包括财务报表、学术论文、扫描发票、市政报告、数据表和新闻通讯文章,所有这些都在 L4 GPU 上处理。Granite-Docling 因其出色的 Markdown 输出(具有正确的标题级别和管道表)而受到关注,而 MinerU 则展示了从图表和图像中提取数据的能力,尽管它默认会静默省略页脚和页面装饰。用户还提到,他们的 Web 应用程序 hexread.com 被用于运行这些基准测试,并提供直接的 PDF 测试,有免费试用。 AI

影响 为开发人员和用户提供了对 PDF 解析模型实际性能差异的见解。

排序理由 用户进行的基准测试,比较了多种开源模型在特定能力上的表现。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

MinerU、Granite-Docling、PaddleOCR-VL 的 PDF 解析能力对比

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/LowerGears ·

    I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vecxhw/i_compared_mineru_granitedocling_and_paddleocrvl/"> <img alt="I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types" src="https://preview.redd.it/pu…