PulseAugur
EN
LIVE 08:23:42

New benchmark evaluates LLM Theory of Mind using asymmetric game mechanics

Researchers have developed Avalon-ToM-Bench, a new benchmark designed to evaluate the fine-grained Theory of Mind (ToM) capabilities of large language models. This benchmark utilizes the asymmetric information mechanics of the game The Resistance: Avalon, focusing on specific aspects of mental-state reasoning rather than end-to-end gameplay. Initial testing on 28 LLMs indicated that models struggle more with social reasoning and expressing inferred mental states than with comprehending game rules or possessing domain knowledge. The findings also suggest that dedicated reasoning training is more effective for improving ToM than simply increasing inference-time deliberation. AI

IMPACT This benchmark could lead to more nuanced evaluations of LLM social reasoning, potentially driving improvements in agent interaction capabilities.

RANK_REASON The cluster contains an academic paper detailing a new benchmark for evaluating AI capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark evaluates LLM Theory of Mind using asymmetric game mechanics

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yen-Shan Chen, Yu Chian Duan, Chih-En Kuo, Jian-Bin Wu, Yun-Nung Chen ·

    Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

    arXiv:2608.09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present …