PulseAugur
实时 15:41:19
English(EN) Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

前沿视觉语言模型在跨任务心智理论方面表现出碎片化 · 跟踪2个来源

一篇新论文评估了九个前沿视觉语言模型(VLMs)在两个心智理论(ToM)基准测试上的表现:Keysar导演任务和Frith-Happé动画三角形任务。研究发现,这些模型表现出碎片化的心智理论能力,在不同任务上的表现不一致。在导演任务上,VLMs经常犯以自我为中心的错误,类似于儿童,尽管推理能力对某些模型有所提升。然而,在动画三角形任务上,模型倾向于低估意图归因,更符合高功能自闭症成年人的特征,而非典型成年人。值得注意的是,没有一个模型在两个测试范式中都展现出成人般的心智理论能力。 AI

影响 揭示了当前前沿VLMs在不同任务上泛化心智理论能力方面的局限性,表明需要更强大和更集成的推理架构。

排序理由 该集群包含一篇详细介绍AI模型研究结果的学术论文。

在 arXiv cs.MA (Multiagent) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

前沿视觉语言模型在跨任务心智理论方面表现出碎片化 · 跟踪2个来源

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Kejia Zhang, Youran Sun, Chugang Yi, Haizhao Yang ·

    前沿视觉语言模型心智理论中的跨任务解离

    arXiv:2608.00261v1 Announce Type: new Abstract: Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nin…

  2. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Haizhao Yang ·

    前沿视觉语言模型心智理论中的跨任务解离

    Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchm…