English(EN)CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
新AI方法提升视频时间定位的准确性和效率
作者PulseAugur 编辑部·[6 个来源]·
研究人员正在开发用于视频时间定位的高级方法,该任务涉及根据自然语言查询在视频中定位事件。几种新方法侧重于通过自动化进化和师生课程耦合来改进引导时间预测的代理工具。其他工作引入了增强该领域视觉语言模型准确性和效率的技术,包括生成更自信的时间预测的方法以及为超长视频协同进化策略与媒体工具。这些进展旨在提高基准测试的性能并降低计算成本。
AI
arXiv:2610.01766v1 Announce Type: new Abstract: Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions …
arXiv cs.AI
TIER_1English(EN)·Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham·
arXiv:2609.40055v1 Announce Type: cross Abstract: On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipe…
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a nov…
arXiv:2609.39883v1 Announce Type: new Abstract: Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explic…
arXiv:2609.40048v1 Announce Type: new Abstract: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool c…
arXiv cs.CV
TIER_1English(EN)·Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou·
arXiv:2609.34581v2 Announce Type: replace Abstract: Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts i…