A new benchmark, WC2026-Agents, has been developed to evaluate the forecasting capabilities of large language models using the 2026 FIFA World Cup as a contamination-free dataset. Four leading models—Claude Opus-4.8, ChatGPT (GPT-5.5), Gemini-3.1 Pro, and Grok—were tasked with predicting match outcomes and making virtual bets. Their performance was compared against the pre-match betting market, revealing that while the models often agreed on predictions, none outperformed the market's Brier score, and a simple market-favorite strategy was more profitable. The benchmark also highlights significant differences in how the models handle decision-making, investment, and self-assessment of errors. AI
IMPACT Highlights limitations in LLM forecasting and decision-making compared to established market baselines, suggesting areas for future model development.
RANK_REASON The item describes a new benchmark and dataset for evaluating LLMs on forecasting tasks, presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →