PulseAugur
实时 09:44:33
English(EN) When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

EcoFrame 框架提升 VLM 在长视频分析中的效率

研究人员开发了 EcoFrame,一个旨在提高视觉语言模型 (VLM) 在长视频理解效率的新框架。与使用静态帧选择或昂贵的多轮推理的先前方法不同,EcoFrame 根据 VLM 的推理反馈来调整其证据收集过程。它采用熵门控预算调度来动态调整帧预算,并通过注意力引导的候选提议来聚焦搜索信息区域。实验表明,EcoFrame 在各种 VLM 上提供了卓越的准确性-效率权衡,在 Video-MMELongVideoBenchMLVU 等基准测试中优于 BOLT 和 A.I.R. 等现有方法。 AI

影响 提高了长视频分析的效率,可能为监控、内容审核和摘要等新应用带来可能。

排序理由 研究论文,详细介绍了 VLM 中高效长视频理解的新框架。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

EcoFrame 框架提升 VLM 在长视频分析中的效率

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen ·

    When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

    arXiv:2608.03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed f…