PulseAugur
中
实时 10:45:17
English(EN) Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

新框架提升AI视频推理效率与准确性

研究人员正在开发用于大型语言模型视频推理的先进方法,旨在提高效率和准确性。Apple的内部化视觉思维(IVT)框架训练模型在内部预测视频的未来状态,与生成中间推理图像相比,降低了推理开销。同时,其他方法侧重于从长视频中提取相关证据,例如PACE,它使用问题派生的因素和候选答案来指导证据获取。新的数据集和基准,如Very Big Video Reasoning (VBVR) Suite和VBVR-Pro,正在被引入,以促进大规模视频推理能力的研究和评估,朝着更强大和可验证的系统迈进。 AI

影响 视频推理的进步可能导致更强大的AI代理,用于需要理解动态视觉环境的任务。

排序理由 多篇研究论文介绍了AI模型视频推理的新框架、数据集和基准。

在 Apple Machine Learning Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

新框架提升AI视频推理效率与准确性

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了AI模型视频推理的新框架、数据集和基准。
Source corroboration
9 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
46 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [9]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    超越视觉CoT:内部化视觉思维以进行主动视频推理

    Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substan…

  2. arXiv cs.LG TIER_1 English(EN) · Baixuan Xu, Yinyui Xu, Tianshi Zheng, Zhaowei Wang, Weiqi Wang, Haochen Shi, Jiayu Liu, Qing Zong, Xiyu Ren, Xinyu Geng, Zhitao He, Yangqiu Song ·

    寻找正确证据:因子引导的粗粒度到细粒度推理用于长视频

    arXiv:2608.26355v1 Announce Type: cross Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alter…

  3. arXiv cs.AI TIER_1 (AF) · Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thadd\"aus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Jiachen Li, Hanwen Xing, Tianqi Zhao, Fengyuan Yu, Wei… ·

    一个非常大的视频推理套件

    arXiv:2602.20159v3 Announce Type: replace-cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyo…

  4. arXiv cs.LG TIER_1 English(EN) · Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehon… ·

    VBVR-Pro:一个可扩展且可验证的原生视觉推理套件

    arXiv:2608.26105v1 Announce Type: cross Abstract: Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem so…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    VBVR-Pro:一个可扩展且可验证的原生视觉推理套件

    VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates.

  6. arXiv cs.AI TIER_1 English(EN) · Daeun Lee, Shoubin Yu, Yue Zhang, Mohit Bansal ·

    VisionCoach:通过视觉感知提示增强基于视频的推理

    arXiv:2603.14659v2 Announce Type: replace-cross Abstract: Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifiable rewards improves accuracy, it still struggles to achieve reliable spatio-tempo…

  7. arXiv cs.AI TIER_1 English(EN) · Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Xuanyu Zheng, Changyi Liu, Tianke Zhang, Haonan Fan, Kaiyu Jiang, Yingxin Li, Jiankang Chen, Xu Wang, Bin Wen, Tingting Gao, Han Li, Jianhua Yin, Yinwei Wei, Xuemeng Song ·

    超越视频思考:为开放世界视频智能体统一视频推理与深度研究

    arXiv:2608.23329v1 Announce Type: cross Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perc…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越视频思考:为开放世界视频智能体统一视频推理与深度研究

    Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step infor…

  9. arXiv cs.CV TIER_1 English(EN) · Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi, Jinsung Yoon, Hyunwoo J. Kim ·

    言语中的理性:视频大模型推理蒸馏的非策略痕迹的语者个性化释义

    arXiv:2608.26684v1 Announce Type: new Abstract: Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generate…