PulseAugur
实时 10:48:10
English(EN) Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Video-DeepResearch 代理在视频问答基准测试上创下新的 SOTA 纪录

研究人员开发了 Video-DeepResearch (Video-DR),这是一种能够处理连续视频流以完成复杂研究任务的多模态代理。该新框架通过采用解耦的感知-探索管道和两阶段训练过程,解决了模态偏差和参数知识泄露问题。在评估中,Video-DeepResearch-35B-A3B 模型在一个具有挑战性的视频问答基准测试上取得了 64.0% 的新最先进准确率,超越了 Claude-4.5-SonnetGPT-5Gemini 2.5 Pro 等领先的专有模型。 AI

影响 为基于视频的人工智能研究代理树立了新的基准,有望推动多模态理解和工具使用的进步。

排序理由 该集群描述了一篇介绍新型多模态代理和基准的研究论文,并与现有模型进行了性能比较。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Video-DeepResearch 代理在视频问答基准测试上创下新的 SOTA 纪录

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao ·

    Video-DeepResearch:迈向下一代多模态深度研究代理

    arXiv:2608.03979v1 Announce Type: cross Abstract: We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluatio…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Video-DeepResearch:迈向下一代多模态深度研究代理

    We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current mode…