PulseAugur
中
实时 15:55:54
English(EN) Video Evidence Indexing: Learning Where to Look from Video Previews for Token-Budgeted Long-Video Question Answering

新框架提升MLLM处理长视频的效率 · 跟踪3个来源

研究人员开发了新的方法来提高多模态大型语言模型(MLLMs)在处理长视频时的效率。FORTE 使用自适应评分和高斯过程来选择与问题相关的帧,平衡预测的相关性与时间覆盖范围。视频证据索引(VEI)采用一种策略,利用视频预览和特权自蒸馏来学习定位相关时刻并规划预算分配。FocusGraph 利用场景图 LLM 选择器和图结构帧选择来识别具身智能体的关键帧,降低推理成本。 AI

影响 这些方法旨在提高 MLLMs 处理长视频内容的效率和有效性,可能为具身智能和视频分析带来新应用。

排序理由 三篇不同的研究论文发表在 arXiv 上,详细介绍了长视频问答的新颖方法。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新框架提升MLLM处理长视频的效率 · 跟踪3个来源

本文如何被排名

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
三篇不同的研究论文发表在 arXiv 上,详细介绍了长视频问答的新颖方法。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.CV TIER_1 English(EN) · Haifeng Huang, Biyin Xu, Chunsheng Xin, Yang Li ·

    FORTE:长视频问答的自适应评分和精确关键帧选择

    arXiv:2610.00573v1 Announce Type: new Abstract: Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typically search within a fixed, unifo…

  2. arXiv cs.CV TIER_1 English(EN) · Haowen Guan, Shengzhi Li, Shichao Pei ·

    视频证据索引:从视频预览中学习在哪里查找以实现令牌预算的长视频问答

    arXiv:2610.00757v1 Announce Type: new Abstract: Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may require broad temporal coverage, but the answer is often supported by only a compac…

  3. arXiv cs.CV TIER_1 English(EN) · Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin ·

    FocusGraph:用于具身长视频问答的图结构帧选择

    arXiv:2603.04349v2 Announce Type: replace Abstract: Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horizon perceptual memories. Multimodal large language models (MLLMs) are increasingly…