PulseAugur
EN
LIVE 14:51:29

Frontier VLMs show fragmented Theory-of-Mind across tasks · 2 sources tracked

A new paper evaluates nine frontier vision-language models (VLMs) on two Theory-of-Mind (ToM) benchmarks: the Keysar Director Task and the Frith-Happé animated triangles. The study found that these models exhibit a fragmented ToM profile, performing inconsistently across different tasks. On the Director Task, the VLMs frequently made egocentric errors, similar to children, though reasoning capabilities improved performance for some. However, on the animated triangles task, the models tended to under-attribute intention, aligning more with a high-functioning autistic adult profile than a typical adult one. Notably, no single model demonstrated adult-like ToM capabilities across both tested paradigms. AI

IMPACT Reveals limitations in current frontier VLMs' ability to generalize Theory-of-Mind capabilities across different tasks, suggesting a need for more robust and integrated reasoning architectures.

RANK_REASON The cluster contains an academic paper detailing research findings on AI models.

Read on arXiv cs.MA (Multiagent) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Frontier VLMs show fragmented Theory-of-Mind across tasks · 2 sources tracked

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Kejia Zhang, Youran Sun, Chugang Yi, Haizhao Yang ·

    Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

    arXiv:2608.00261v1 Announce Type: new Abstract: Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nin…

  2. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Haizhao Yang ·

    Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

    Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchm…