PulseAugur
中
实时 23:33:47
English(EN) The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

本地AI模型:文档提取的准确性与能耗权衡

一篇新的研究论文探讨了小型本地语言模型在处理敏感文档信息提取时,准确性与能耗之间的权衡。该研究遵循本地处理限制,发现批量处理请求在不牺牲准确性的前提下,能显著降低能耗。信息提取的最佳方法取决于文档的布局,视觉语言模型在视觉丰富的文档上表现更好,而带有解析器的纯文本模型在纯文本上表现更佳。 AI

影响 为节能、合规的本地信息提取提供了指导,影响敏感数据的部署策略。

排序理由 arXiv上发表的研究论文,详细介绍了本地AI模型的准确性-能耗权衡。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

本地AI模型:文档提取的准确性与能耗权衡

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst ·

    正确的信息提取管道取决于文档:小型本地模型的准确性-能耗权衡

    arXiv:2609.31341v1 Announce Type: new Abstract: Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud …