PulseAugur
EN
LIVE 14:08:07

New EAST benchmark reveals LLM theory of mind gaps

Researchers have developed a new evaluation method called the Epistemic Asymmetry Schelling Task (EAST) to assess Theory of Mind (ToM) in large language models (LLMs). Unlike traditional tests like the Sally-Anne task, EAST uses a two-player dialogue game to measure robust social reasoning and coordination abilities. The study found that while frontier models show some success, many LLMs struggle with epistemic tracking, often confusing private knowledge with mutual knowledge, indicating a significant gap in functional social reasoning. AI

IMPACT Highlights critical gaps in LLM social reasoning, guiding future development towards more robust AI.

RANK_REASON The cluster describes a new academic paper introducing a novel evaluation method for LLMs.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New EAST benchmark reveals LLM theory of mind gaps

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Roberta Rocca, Sami Boukortt, Geoff Keeling, Winnie Street ·

    Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

    arXiv:2607.11363v1 Announce Type: cross Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obvi…

  2. arXiv cs.AI TIER_1 English(EN) · Winnie Street ·

    Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

    Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in way…