PulseAugur
实时 14:00:15
English(EN) How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

视觉语言模型在鲁棒性、因果推理和视觉搜索方面接受测试

研究人员正在从多个维度调查视觉语言模型(VLM)的鲁棒性和推理能力。一项研究引入了OCR-Robust,这是一个用于评估VLM在光学字符识别任务中对视觉扰动的韧性的基准,揭示了图表和表格等结构性元素特别脆弱。另一篇论文探讨了VLM在因果顺序推理方面的挣扎,发现它们尽管在物体识别方面表现出色,但由于训练数据中缺乏明确的因果表达,因此表现不佳。此外,一项研究检查了VLM执行视觉搜索任务的情况,将其“推理令牌”使用与人类反应时间进行比较,并指出了它们搜索策略的相似性和差异性。 AI

影响 这些研究突显了当前视觉语言模型的一些关键局限性,特别是在对视觉损坏的鲁棒性、因果推理和类似人类的视觉搜索方面,为未来的研究和开发提供了指导。

排序理由 该集群由arXiv上发表的多篇学术论文组成,详细介绍了视觉语言模型的研究。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

视觉语言模型在鲁棒性、因果推理和视觉搜索方面接受测试

报道来源 [6]

  1. arXiv cs.CL TIER_1 English(EN) · Yuxing Cheng, Yuan Wu, Yi Chang ·

    OCR-Reasoning 的鲁棒性如何?评估视觉语言模型在视觉扰动下的 OCR-Reasoning 鲁棒性

    arXiv:2606.26041v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood…

  2. arXiv cs.CL TIER_1 English(EN) · Yi Chang ·

    OCR-Reasoning 的鲁棒性如何?评估视觉语言模型在视觉扰动下的 OCR-Reasoning 鲁棒性

    Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where vi…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    OCR-Reasoning 的鲁棒性如何?评估视觉语言模型在视觉扰动下的 OCR-Reasoning 鲁棒性

    Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where vi…

  4. arXiv cs.AI TIER_1 English(EN) · Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu ·

    当语言覆盖视觉:视觉-语言模型中的过度对齐与几何消偏

    arXiv:2605.08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input. We investigate…

  5. arXiv cs.CL TIER_1 English(EN) · Zhaotian Weng, Haoxuan Li, Xin Eric Wang, Kuan-Hao Huang, Jieyu Zhao ·

    视觉-语言模型缺少什么?探究其因果顺序推理的困境

    arXiv:2506.00869v3 Announce Type: replace Abstract: Despite the impressive performance of vision-language models (VLMs) on downstream tasks, their ability to understand and reason about causal relationships in visual inputs remains unclear. Robust causal reasoning is fundamental …

  6. arXiv cs.CV TIER_1 English(EN) · Farahnaz Wick ·

    视觉语言模型会像人类一样搜索吗?经典视觉搜索范式中的推理令牌作为反应时模拟

    arXiv:2606.25066v1 Announce Type: cross Abstract: Visual search has been one of the most productive paradigms in the study of visual attention: the way reaction time scales with the number of items distinguishes parallel, "pop-out" search from serial, attention-demanding search. …