PulseAugur
实时 12:08:50
English(EN) Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

VideoLLMs 表现出“事件包”行为,臆想时间联系

一篇新发表在 arXiv 上的研究介绍了 DistractionBench,一个旨在测试视频大语言模型 (VideoLLMs) 时间理解能力的框架。研究人员发现,这些模型经常表现出“事件包”行为,这意味着它们将视频视为一系列不相关的事件集合,而不是一个连贯、有时间结构序列。这会导致显著的臆想,模型会错误地将插入片段(如广告)中的动作归因于主视频内容中的主体。该研究评估了 11 个流行的 VideoLLMs,所有模型都显示出这一缺陷,突显了未来模型需要改进时间关联机制。 AI

影响 强调了当前 VideoLLMs 的一个关键缺陷,表明需要改进时间关联和主体-事件关联,以实现更可靠的视频理解。

排序理由 在 arXiv 上发表的研究论文,详细介绍了一个新的评估框架和关于 VideoLLMs 的发现。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

VideoLLMs 表现出“事件包”行为,臆想时间联系

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Oscar Chew, Serhii Honcharenko, Qian-Hui Chen, Patricia Lu, Dishant Zaveri, Khoa D. Doan, Kuan-Hao Huang ·

    弹出式干扰揭示视频大语言模型中的事件包行为

    arXiv:2605.27101v1 Announce Type: cross Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to…

  2. arXiv cs.CL TIER_1 English(EN) · Kuan-Hao Huang ·

    弹出式干扰揭示视频大语言模型中的事件包行为

    A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subj…