PulseAugur
实时 05:01:06
English(EN) When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

视觉语言模型在历史文献OCR中表现出细微的幻觉

一篇新的研究论文分析了视觉语言模型(VLM)在转录历史文献方面的性能。研究发现,尽管它们在字符错误率(CER)和单词错误率(WER)等标准指标上优于传统的OCR系统,但它们表现出细微但关键的故障模式。这些模式包括生成虚假内容和进行语义替换,这些替换在不显著影响CER/WER的情况下改变了含义,尤其影响了命名实体。该研究强调,需要采用评估方法来评估语义可靠性,而不仅仅是字符准确性,以满足档案转录的需求。 AI

影响 突出了VLM在历史文献转录中关键的语义可靠性差距,需要新的评估框架。

排序理由 一篇在arXiv上发表的研究论文,详细介绍了视觉语言模型在OCR方面的局限性。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

视觉语言模型在历史文献OCR中表现出细微的幻觉

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Marina Gardella (CB), Camilo Mari{\~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{\'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong) ·

    当低CER不足以说明问题:对历史乌拉圭文件上视觉语言OCR系统幻觉的分析

    arXiv:2607.24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art …