PulseAugur
实时 10:08:53
English(EN) V-ICAL Bench: Evaluating Video In-Context Learning for Multimodal Agents in Interactive Environments

新的V-ICAL基准揭示了多模态智能体基于视频学习的显著局限性

引入了一个名为V-ICAL的新基准,用于评估多模态智能体在交互式环境中从视频演示中学习的能力。该基准包含37个环境中的342个任务,评估智能体将视频示例转化为可执行策略以及适应新情况的能力。包括Seed-2.1-Pro、Gemini-3.1-Pro和GPT-5.6在内的当前最先进的智能体显示出显著的局限性,表现最佳的智能体得分仅为54.4,表明这些智能体在基于视频的上下文学习能力方面存在重大差距。 AI

影响 突出了多模态智能体能力方面的重大差距,有必要在基于视频的上下文学习方面取得进展。

排序理由 该集群描述了一个新的学术基准和对现有模型的评估,符合研究类别。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的V-ICAL基准揭示了多模态智能体基于视频学习的显著局限性

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个新的学术基准和对现有模型的评估,符合研究类别。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Ziqian Fan, Shibo Xu, Junjie Li, Xiangyu Zhao, Shengyuan Ding, Yifan Yang, Zhenjie Yang, Haodong Duan, Yue Zhou, Zhihang Zhong, Xue Yang ·

    V-ICAL 评估:在交互式环境中评估多模态智能体的视频上下文学习能力

    arXiv:2609.15683v1 Announce Type: new Abstract: While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particularly regarding video demonstrations in interactive environments. For multimodal …