PulseAugur
EN
LIVE 11:00:46

New benchmark tests VLMs' strategic decision-making in soccer

Researchers have developed SportD, a new benchmark designed to test the strategic decision-making capabilities of vision-language models (VLMs) in physical environments. The benchmark uses on-ball decisions from the 2022 FIFA World Cup, evaluating VLMs' ability to choose between actions like shooting or passing based on estimated possession value. Current frontier VLMs perform below professional players, selecting optimal actions less frequently and exhibiting a preference for safer, less progressive plays, indicating a need for improved strategic reasoning in these models. AI

IMPACT This benchmark could drive the development of VLMs with improved strategic reasoning for real-world applications beyond simple image interpretation.

RANK_REASON The cluster describes a new benchmark and research paper published on arXiv, evaluating AI model capabilities.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New benchmark tests VLMs' strategic decision-making in soccer

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Jasin Cekinmez, Addison J. Wu, Haotian Xia, Akshaya Bharadhwaj, Anay Putty, Anirudh Ravishankar, Jaewoong Lee, Jinglin Xiao, Kyumin Andrew Shim, Mishika Ahuja, Nisarga Patil, Leo Liu, Zhuohan Liu, Weining Shen ·

    SportD: Can VLMs Physically Strategize?

    arXiv:2607.14616v1 Announce Type: new Abstract: Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where…

  2. arXiv cs.CV TIER_1 English(EN) · Weining Shen ·

    SportD: Can VLMs Physically Strategize?

    Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where models observe the seconds preceding an on-ball…