English(EN)Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
新框架提升AI视频推理效率与准确性
作者PulseAugur 编辑部·[9 个来源]·
研究人员正在开发用于大型语言模型视频推理的先进方法,旨在提高效率和准确性。Apple的内部化视觉思维(IVT)框架训练模型在内部预测视频的未来状态,与生成中间推理图像相比,降低了推理开销。同时,其他方法侧重于从长视频中提取相关证据,例如PACE,它使用问题派生的因素和候选答案来指导证据获取。新的数据集和基准,如Very Big Video Reasoning (VBVR) Suite和VBVR-Pro,正在被引入,以促进大规模视频推理能力的研究和评估,朝着更强大和可验证的系统迈进。
AI
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substan…
arXiv:2608.26355v1 Announce Type: cross Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alter…
arXiv:2602.20159v3 Announce Type: replace-cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyo…
arXiv cs.LG
TIER_1English(EN)·Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehon…·
arXiv:2608.26105v1 Announce Type: cross Abstract: Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem so…
VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates.
arXiv:2603.14659v2 Announce Type: replace-cross Abstract: Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifiable rewards improves accuracy, it still struggles to achieve reliable spatio-tempo…
arXiv:2608.23329v1 Announce Type: cross Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perc…
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step infor…
arXiv cs.CV
TIER_1English(EN)·Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi, Jinsung Yoon, Hyunwoo J. Kim·
arXiv:2608.26684v1 Announce Type: new Abstract: Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generate…