A new research paper published on arXiv evaluates six frontier language models on a communication efficiency task called the log(N)-Questions game. The game involves two instances of the same model, one acting as a questioner and the other as an answerer, to identify a target document from a set of N Wikipedia abstracts using only log(N) yes/no questions. Claude Opus-5 performed significantly worse than GLM 5.3, GPT 5.6 "Sol", Grok 4.6, Gemini 3.8 Flash, and Kimi K3, with win rates declining as the document set size increased. The study found that information per question correlates with win rate, and models that partition documents by title performed better. AI
IMPACT Highlights differences in communication efficiency and strategic document partitioning among leading LLMs, potentially influencing future model development.
RANK_REASON Research paper evaluating frontier language models on a novel task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →