PulseAugur
实时 08:45:17
English(EN) Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

Claude Opus-5 在维基百科摘要通信游戏中落后于前沿模型 · arXiv 研究

一篇新发表在 arXiv 上的研究论文评估了六种前沿语言模型在一项名为 log(N)-Questions 游戏(通信效率任务)中的表现。该游戏涉及同一模型的两个实例,一个充当提问者,另一个充当回答者,仅通过 log(N) 个是/否问题来从 N 个维基百科摘要的集合中识别目标文档。Claude Opus-5 的表现明显逊于 GLM 5.3、GPT 5.6 "Sol"、Grok 4.6Gemini 3.8 FlashKimi K3,并且随着文档集大小的增加,其胜率下降。研究发现,每个问题的平均信息量与胜率相关,而通过标题对文档进行分区的模型表现更好。 AI

影响 突出了领先 LLM 在通信效率和策略性文档分区方面的差异,可能影响未来的模型开发。

排序理由 评估前沿语言模型在一项新任务上的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Claude Opus-5 在维基百科摘要通信游戏中落后于前沿模型 · arXiv 研究

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
评估前沿语言模型在一项新任务上的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Peter Potash ·

    在维基百科摘要上玩 log(N) 个问题:配对前沿模型间的通信效率

    arXiv:2609.19113v1 Announce Type: new Abstract: We evaluate six frontier language models on the two-agent $\log(N)$-Questions game. A questioner sees $N$ Wikipedia lead paragraphs and must identify a secretly chosen target using exactly $\log_2 N$ yes/no questions. An answerer se…