PulseAugur
实时 11:59:07
English(EN) SportD: Can VLMs Physically Strategize?

新基准测试视觉语言模型在足球中的策略决策能力

研究人员开发了SportD,这是一个旨在测试视觉语言模型(VLMs)在物理环境中的策略决策能力的新基准。该基准使用2022年FIFA世界杯的带球决策,评估VLMs根据估计的控球价值在射门或传球等动作之间进行选择的能力。目前最前沿的VLMs表现不如职业球员,选择最优动作的频率较低,并且倾向于选择更安全、进步性较差的打法,这表明这些模型在策略推理方面仍需改进。 AI

影响 该基准有望推动VLMs在超越简单图像解释的现实世界应用中,具备更强的策略推理能力。

排序理由 该集群描述了一个新基准和在arXiv上发表的研究论文,评估了AI模型的能力。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新基准测试视觉语言模型在足球中的策略决策能力

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen ·

    SportD:VLMs能否进行物理策略制定?

    arXiv:2607.14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where…

  2. arXiv cs.CV TIER_1 English(EN) · Weining Shen ·

    SportD:VLMs能否进行物理策略制定?

    Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where models observe the seconds preceding an on-ball…