PulseAugur
EN
LIVE 09:43:51

New ACT-Eval framework reveals pervasive hallucinations in LLM chess commentary

A new evaluation framework called ACT-Eval has been developed to assess the factual accuracy and conceptual coverage of large language model (LLM) chess commentary. This framework decomposes commentary into atomic claims and uses engine-supported tools and expert references to verify correctness. Evaluations show that even advanced models like GPT-5.4 hallucinate incorrect sub-claims 22.0% of the time without tool augmentation, while smaller open-weight models exceed 40%. While tool augmentation significantly improves factual correctness and move-quality assessment, the models still struggle to cover expert strategic and tactical ideas comprehensively. AI

IMPACT Highlights the ongoing challenge of factual accuracy in LLM outputs, even in specialized domains, and the need for robust evaluation tools.

RANK_REASON The cluster contains an academic paper detailing a new evaluation framework and benchmark for LLM commentary. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New ACT-Eval framework reveals pervasive hallucinations in LLM chess commentary

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh, Pramod Viswanath ·

    Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

    arXiv:2608.04240v1 Announce Type: cross Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to exper…