PulseAugur
EN
LIVE 00:24:59

New EAST benchmark reveals LLM theory of mind gaps

Researchers have developed a new evaluation method called the Epistemic Asymmetry Schelling Task (EAST) to assess Theory of Mind (ToM) in large language models (LLMs). Unlike traditional tests like the Sally-Anne task, EAST uses a two-player dialogue game to measure robust social reasoning and coordination abilities. The study found that while frontier models show some success, many LLMs struggle with epistemic tracking, often confusing private knowledge with mutual knowledge, indicating a significant gap in functional social reasoning. AI

IMPACT Highlights critical gaps in LLM social reasoning, guiding future development towards more robust AI.

RANK_REASON The cluster describes a new academic paper introducing a novel evaluation method for LLMs.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New EAST benchmark reveals LLM theory of mind gaps

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster describes a new academic paper introducing a novel evaluation method for LLMs.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
70 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Roberta Rocca, Sami Boukortt, Geoff Keeling, Winnie Street ·

    Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

    arXiv:2607.11363v1 Announce Type: cross Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obvi…

  2. arXiv cs.AI TIER_1 English(EN) · Winnie Street ·

    Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

    Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in way…