PulseAugur
实时 02:22:28
English(EN) SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

新的SER方法通过语义证据奖励增强视频MLLM推理 · 跟踪4个来源

研究人员开发了一种名为语义证据奖励(SER)的新方法,以提高视频多模态大语言模型(Video MLLMs)的时空推理能力。现有模型在细粒度推理方面常常遇到困难,有时会使用不相关的帧或对象来回答问题。SER通过将证据定位重构为验证任务来解决这个问题,使用一个裁判VLM来评估生成证据的相关性和定位质量,从而减少对密集标注的需求。这种方法通过在V-STAR基准上提高3.0个点来增强答案准确性和证据定位,如所证明的。 AI

影响 增强了Video MLLM的准确性和证据定位,可能减少视频问答任务对大量标注的依赖。

排序理由 该集群包含多篇arXiv论文,详细介绍了一种用于改进Video MLLMs的新研究方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

新的SER方法通过语义证据奖励增强视频MLLM推理 · 跟踪4个来源

报道来源 [7]

  1. arXiv cs.AI TIER_1 English(EN) · Chaohong Guo, Xun Mo, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long ·

    TAR:视频时间定位的临时锚点约束推理

    arXiv:2508.07683v2 Announce Type: replace-cross Abstract: Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforcement Learning to generate Chains-of-Thoug…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    SER:通过语义证据奖励学习视频推理的接地

    Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising direction, existing RL frameworks typically rely on geo…

  3. arXiv cs.CV TIER_1 English(EN) · Zhengbo Zhang, Mark He Huang, Zhigang Tu, Ming-Hsuan Yang ·

    DART:一种用于零样本视频时序定位的难度自适应路由

    arXiv:2607.00672v1 Announce Type: new Abstract: Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple events but …

  4. arXiv cs.CV TIER_1 English(EN) · Ming-Hsuan Yang ·

    DART:用于零样本视频时序定位的难度自适应路由

    Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple events but struggles with complex multi-stage queries that …

  5. arXiv cs.CV TIER_1 English(EN) · Jiaqi Li, Shuntian Zheng, Yixian Shen, Jia-Hong Huang, Xiaoman Lu, Minzhe Ni, Yu Guan ·

    保持证据链:无训练的视频时序定位令牌剪枝的语义证据分配

    arXiv:2603.05663v3 Announce Type: replace Abstract: Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long, untrimmed videos, making video-language-model prohibitively expensive. While recent training-free token pruning has shown succes…

  6. arXiv cs.CV TIER_1 English(EN) · Sheng Xia, Zhengqin Lai, Tianxiang Jiang, Kanghui Tian, Shoujun Zhou, Bin Li, Yi Wang ·

    SER:通过语义证据奖励学习视频推理的接地

    arXiv:2606.24726v1 Announce Type: new Abstract: Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising directi…

  7. arXiv cs.CV TIER_1 English(EN) · Yi Wang ·

    SER:通过语义证据奖励学习视频推理的接地

    Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising direction, existing RL frameworks typically rely on geo…