PulseAugur
中
实时 23:55:38
English(EN) CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

新AI方法提升视频时间定位的准确性和效率

研究人员正在开发用于视频时间定位的高级方法,该任务涉及根据自然语言查询在视频中定位事件。几种新方法侧重于通过自动化进化和师生课程耦合来改进引导时间预测的代理工具。其他工作引入了增强该领域视觉语言模型准确性和效率的技术,包括生成更自信的时间预测的方法以及为超长视频协同进化策略与媒体工具。这些进展旨在提高基准测试的性能并降低计算成本。 AI

影响 这些进展可能带来更高效、更准确的视频分析工具,改进视频搜索和内容审查等应用。

排序理由 多篇研究论文发表在arXiv上,详细介绍了视频时间定位的新技术。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新AI方法提升视频时间定位的准确性和效率

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文发表在arXiv上,详细介绍了视频时间定位的新技术。
Source corroboration
6 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [6]

  1. arXiv cs.AI TIER_1 Norsk(NO) · Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li ·

    VideoEvolve:用于视频时间定位的演进式智能体集线器

    arXiv:2610.01766v1 Announce Type: new Abstract: Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions …

  2. arXiv cs.AI TIER_1 English(EN) · Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham ·

    数据更少,时机更佳:用于VLM on-policy蒸馏和时序视频定位的学生-课程耦合

    arXiv:2609.40055v1 Announce Type: cross Abstract: On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipe…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoEvoWhen: 策略-工具协同进化用于超长视频时间定位

    Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a nov…

  4. arXiv cs.CV TIER_1 English(EN) · Jinhao Chen, Benlei Cui, Ruijian Jia, Ziheng Wang, Tianyu Wo, Pengfei Sun, Longtao Huang, Hui Xue, Yitong Yang, Haiwen Hong ·

    自信地进行地面定位:可控生成视频时间地面定位

    arXiv:2609.39883v1 Announce Type: new Abstract: Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explic…

  5. arXiv cs.CV TIER_1 English(EN) · Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen ·

    CoEvoWhen:用于超长视频时间定位的策略-工具协同进化

    arXiv:2609.40048v1 Announce Type: new Abstract: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool c…

  6. arXiv cs.CV TIER_1 English(EN) · Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou ·

    面向时序视频定位的逆事实注意力策略蒸馏

    arXiv:2609.34581v2 Announce Type: replace Abstract: Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts i…