PulseAugur
实时 10:00:58
English(EN) MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

新的MotionBlind基准揭示视频大语言模型在运动理解方面存在不足

一个名为MotionBlind的新基准已被开发出来,用于测试视频大语言模型(Video-LLMs)的运动理解能力。研究人员发现,大多数开源的Video-LLMs表现不佳,即使在呈现清晰的视觉数据时,也常常无法区分不同的运动速度或方向。虽然Gemini-3.1 Pro有所改进,但在准确评估速度方面仍有困难,这表明当前的Video-LLMs在需要真正理解运动的任务中尚不可靠。 AI

影响 突出了当前Video-LLMs的关键局限性,表明它们尚未适用于需要细微运动感知的应用。

排序理由 介绍用于评估Video-LLMs新基准的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的MotionBlind基准揭示视频大语言模型在运动理解方面存在不足

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
介绍用于评估Video-LLMs新基准的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas ·

    MotionBlind:探究视频大模型运动理解的幻觉

    arXiv:2609.09528v1 Announce Type: new Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, which way it travels, how hard it is pushed. We show the…