PulseAugur
中
实时 09:44:10

全模态大语言模型难以检测感官与文本信息冲突

研究人员在全模态大语言模型中发现了一个“表征-行动鸿沟”,即模型能够编码文本声明与其感官输入之间的不匹配,但在其输出中却无法基于此信息采取行动。研究人员使用电影片段开发了一个新的基准测试IMA VB,以测试这种冲突检测能力。在八个开源模型和Gemini 3.1 Pro上进行的研究发现,模型要么低估拒绝错误声明,要么过度拒绝,从而影响理解准确性。这种鸿沟在音频方面比在视觉方面更明显,并且对提示具有抵抗力,尽管一种探针引导的logit调整技术在改善拒绝行为方面显示出希望。 AI

影响 凸显了大语言模型基础中的一个关键鸿沟,表明当前模型可能会误解或忽略冲突的感官数据,影响其作为代理的可靠性。

排序理由 该集群包含一篇详细介绍新基准和全模态大语言模型研究结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

全模态大语言模型难以检测感官与文本信息冲突

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍新基准和全模态大语言模型研究结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Trung Nguyen Quang, Yiming Gao, Fanyi Pu, Kaichen Zhang, Shuo Sun, Ziwei Liu ·

    Senses Wide Shut:全模态大语言模型中的表征-行动鸿沟

    arXiv:2605.13737v2 Announce Type: replace Abstract: When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-gr…