A new paper evaluates nine frontier vision-language models (VLMs) on two Theory-of-Mind (ToM) benchmarks: the Keysar Director Task and the Frith-Happé animated triangles. The study found that these models exhibit a fragmented ToM profile, performing inconsistently across different tasks. On the Director Task, the VLMs frequently made egocentric errors, similar to children, though reasoning capabilities improved performance for some. However, on the animated triangles task, the models tended to under-attribute intention, aligning more with a high-functioning autistic adult profile than a typical adult one. Notably, no single model demonstrated adult-like ToM capabilities across both tested paradigms. AI
IMPACT Reveals limitations in current frontier VLMs' ability to generalize Theory-of-Mind capabilities across different tasks, suggesting a need for more robust and integrated reasoning architectures.
RANK_REASON The cluster contains an academic paper detailing research findings on AI models.
Read on arXiv cs.MA (Multiagent) →
- Castelli rubric
- Chad
- Director Task
- Frith-Happé animated triangles
- Goal-directed Self-tracking in the Management of Chronic Health Conditions
- HF-ASD
- Keysar Director Task
- Raheem Jarbo
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →