PulseAugur
中
实时 11:12:19
English(EN) FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

LLM在2026年国际足联世界杯预测基准中未能击败博彩市场

一项新的基准测试WC2026-Agents已被开发出来,用于评估大型语言模型使用2026年国际足联世界杯作为无污染数据集的预测能力。四个领先的模型——Claude Opus-4.8、ChatGPT (GPT-5.5)、Gemini-3.1 Pro和Grok——被要求预测比赛结果并进行虚拟投注。它们的表现与赛前博彩市场进行了比较,结果显示,尽管模型在预测上经常达成一致,但没有一个模型的Brier分数优于市场,并且简单的市场热门策略更具盈利性。该基准还突显了模型在处理决策、投资和错误自我评估方面存在的显著差异。 AI

影响 强调了与既有市场基准相比,LLM在预测和决策方面的局限性,并为未来的模型开发指明了方向。

排序理由 该项目在一篇学术论文中描述了一个用于评估LLM预测任务的新基准和数据集。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM在2026年国际足联世界杯预测基准中未能击败博彩市场

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目在一篇学术论文中描述了一个用于评估LLM预测任务的新基准和数据集。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
79 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jiacheng Ding, Cong Guo, Jason Xu ·

    2026年FIFA世界杯作为LLM预测代理的无污染基准:四种模型、一家博彩公司和104场比赛

    arXiv:2607.17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier mo…