PulseAugur
实时 14:02:57
English(EN) Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

新的EAST基准揭示LLM心理理论的差距

研究人员开发了一种名为认知不对称谢林任务(EAST)的新评估方法,用于评估大型语言模型(LLM)的心理理论(ToM)。与Sally-Anne任务等传统测试不同,EAST使用一个双人对话游戏来衡量强大的社会推理和协调能力。研究发现,虽然前沿模型取得了一些成功,但许多LLM在认知追踪方面存在困难,常常将私有知识与相互知识混淆,这表明在功能性社会推理方面存在显著差距。 AI

影响 强调了LLM社会推理中的关键差距,指导未来朝着更强大的AI发展。

排序理由 该集群描述了一篇介绍LLM新评估方法的学术论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的EAST基准揭示LLM心理理论的差距

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Roberta Rocca, Sami Boukortt, Geoff Keeling, Winnie Street ·

    超越Sally-Anne测试:使用认知性谢林点评估LLM的心理理论

    arXiv:2607.11363v1 Announce Type: cross Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obvi…

  2. arXiv cs.AI TIER_1 English(EN) · Winnie Street ·

    超越Sally-Anne测试:使用认知谢林点评估LLM的心理理论

    Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in way…