English(EN)SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
新的SER方法通过语义证据奖励增强视频MLLM推理 · 跟踪4个来源
作者PulseAugur 编辑部·[7 个来源]·
研究人员开发了一种名为语义证据奖励(SER)的新方法,以提高视频多模态大语言模型(Video MLLMs)的时空推理能力。现有模型在细粒度推理方面常常遇到困难,有时会使用不相关的帧或对象来回答问题。SER通过将证据定位重构为验证任务来解决这个问题,使用一个裁判VLM来评估生成证据的相关性和定位质量,从而减少对密集标注的需求。这种方法通过在V-STAR基准上提高3.0个点来增强答案准确性和证据定位,如所证明的。
AI
arXiv:2508.07683v2 Announce Type: replace-cross Abstract: Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforcement Learning to generate Chains-of-Thoug…
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising direction, existing RL frameworks typically rely on geo…
arXiv cs.CV
TIER_1English(EN)·Zhengbo Zhang, Mark He Huang, Zhigang Tu, Ming-Hsuan Yang·
arXiv:2607.00672v1 Announce Type: new Abstract: Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple events but …
Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple events but struggles with complex multi-stage queries that …
arXiv:2603.05663v3 Announce Type: replace Abstract: Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long, untrimmed videos, making video-language-model prohibitively expensive. While recent training-free token pruning has shown succes…
arXiv cs.CV
TIER_1English(EN)·Sheng Xia, Zhengqin Lai, Tianxiang Jiang, Kanghui Tian, Shoujun Zhou, Bin Li, Yi Wang·
arXiv:2606.24726v1 Announce Type: new Abstract: Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising directi…
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising direction, existing RL frameworks typically rely on geo…