PulseAugur
EN
LIVE 07:25:21

New benchmark reveals vision-language models fail to translate theory of mind into action

Researchers have developed MOSAIC, a new benchmark designed to measure how well vision-language models (VLMs) can translate theory of mind (ToM) reasoning into coordinated social actions. In evaluations of 13 models, including 11 VLMs, it was found that these models struggled to produce behaviors consistent with ToM constraints and showed no significant behavioral change when explicit ToM constraints were imposed. The analysis indicated bottlenecks in generating coherent nonverbal signals and in interpreting and reacting to other agents' behaviors. A model named PCM-LLM, which incorporated an explicit ToM module, succeeded across all conditions, suggesting that integrating belief-action coupling is crucial for these tasks. AI

IMPACT This research highlights critical limitations in current AI's ability to perform nuanced social interactions, suggesting a need for architectural improvements in belief-action coupling for more sophisticated AI agents.

RANK_REASON The cluster contains a research paper detailing a new benchmark and evaluation of AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals vision-language models fail to translate theory of mind into action

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf ·

    Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models

    arXiv:2608.20975v1 Announce Type: new Abstract: Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoni…