Researchers have developed a new evaluation method called the Epistemic Asymmetry Schelling Task (EAST) to assess Theory of Mind (ToM) in large language models (LLMs). Unlike traditional tests like the Sally-Anne task, EAST uses a two-player dialogue game to measure robust social reasoning and coordination abilities. The study found that while frontier models show some success, many LLMs struggle with epistemic tracking, often confusing private knowledge with mutual knowledge, indicating a significant gap in functional social reasoning. AI
IMPACT Highlights critical gaps in LLM social reasoning, guiding future development towards more robust AI.
RANK_REASON The cluster describes a new academic paper introducing a novel evaluation method for LLMs.
- Epistemic Asymmetry Schelling Task
- large-language models
- LLM-LLM dyads
- Sally-Anne task
- Schelling points
- theory of mind
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →