PulseAugur
中
实时 05:59:31
English(EN) Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models

新的基准和架构推动了 MLLMs 中的长视频理解

研究人员正在开发新的方法来改进多模态大语言模型 (MLLMs) 理解长视频的方式。一种方法 MoTE 使用任务专家混合 (Mixture of Task Experts) 将计算路由到特定任务模块,从而提高视频语言任务的准确性和效率。另一个重点是优化模型如何从长视频中提取相关信息,例如 VideoHarness-RSI 通过递归搜索有效的上下文构建程序,以及 TRACE 专注于证据支持的答案。此外,还推出了 Video-IFBench 作为基准,用于评估 MLLMs 在视频理解场景中遵循各种指令的能力,突出了当前在约束遵守方面面临的挑战。 AI

影响 长视频理解能力的进步可能为更复杂的 AI 代理和复杂视觉数据分析工具带来可能。

排序理由 多篇研究论文介绍了用于长视频理解的新架构、基准和方法。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

新的基准和架构推动了 MLLMs 中的长视频理解

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了用于长视频理解的新架构、基准和方法。
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
47 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [13]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    LongVU-TTT:用于长视频理解中视觉重采样因果测试时训练

    Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates between the vision encoder and the LLM. Its grou…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    观察视角至关重要:长视频理解的策略内自蒸馏

    Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the full video. However, as videos become longer, the full-video context inevitably contains more quest…

  3. arXiv cs.AI TIER_1 English(EN) · Guoyang Xu, Hao Chen ·

    VideoHarness-RSI:用于长视频理解的递归式Harness自改进,配合冻结的视觉语言模型

    arXiv:2608.24302v1 Announce Type: new Abstract: Long-video understanding depends critically on how a limited model context is constructed from a much longer video. Existing approaches improve this process through compression, retrieval, memory, and agentic evidence acquisition, b…

  4. arXiv cs.LG TIER_1 English(EN) · Muhammad Asad Ali, Umar Khan, Nadia Robertini, Didier Stricker ·

    MoTE:用于多任务视频理解的混合任务专家

    arXiv:2608.24763v1 Announce Type: cross Abstract: Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks ac…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    Video-IFBench:评估多模态大模型在视频理解场景下的指令遵循能力

    A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints.

  6. arXiv cs.CL TIER_1 English(EN) · Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu ·

    TRACE:用于长时视频理解的具有锚定和收敛证据的时间检索

    arXiv:2608.22516v1 Announce Type: cross Abstract: A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness or predicted evidence intervals, but the frames a m…

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    MoTE:用于多任务视频理解的混合任务专家

    MoTE replaces dense decoder feed-forward networks with task-specific experts routed by sample-level task labels, improving multi-task video-language accuracy with sparse, interpretable computation.

  8. arXiv cs.CV TIER_1 English(EN) · Ziyue Wang, Shiqi Huang, Weiwen Xu, Bihan Wen, Xudong Jiang ·

    Video-OPSD:利用特权视觉证据进行视频大语言模型上的策略自蒸馏

    arXiv:2608.27065v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has recently emerged as an effective post-training paradigm that improves policy optimization through dense token-level supervision from a privileged self-teacher. Despite its promise, OPSD remains…

  9. arXiv cs.CV TIER_1 English(EN) · Kaishen Wang, Dongdi Zhao, Yijun Liang, Dingqiang Ye, Ruibo Chen, Heng Huang, Di Fu ·

    观察视角至关重要:长视频理解的策略内自蒸馏

    arXiv:2608.25356v1 Announce Type: new Abstract: Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the full video. However, as videos become longer, the f…

  10. arXiv cs.CV TIER_1 English(EN) · Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao ·

    Video-IFBench:评估多模态大模型在视频理解场景下的指令遵循能力

    arXiv:2608.25529v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not o…

  11. arXiv cs.CV TIER_1 English(EN) · Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny ·

    LongVU-TTT:用于长视频理解中视觉重采样的因果测试时训练

    arXiv:2608.25729v1 Announce Type: new Abstract: Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates …

  12. arXiv cs.CV TIER_1 English(EN) · Tianyue Wang, Xuying Wu, Yuxiang Ma, Ruiming Liang, Jiaxuan Kang, Yanchao Hao, Zheng Wei, Leigang Qu, Haiyun Guo, Jinqiao Wang ·

    路由优先于查找:查询自适应证据获取用于长视频理解

    arXiv:2608.20805v1 Announce Type: new Abstract: Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-perception methods outperform query-agnostic pipelines…

  13. arXiv cs.CV TIER_1 English(EN) · Beibei Zhang, Chao Xu, Jun Lan, Zongyi Li, Lai Wei, Huijia Zhu, Tongwei Ren ·

    通过高效的片段到视频监督增强长视频理解的本地化推理

    arXiv:2608.20814v1 Announce Type: new Abstract: Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise in complex and lengthy contexts can obscure localized…