PulseAugur
EN
LIVE 20:06:59

FutureSim benchmark tests AI forecasting with historical data

Researchers from the Max Planck Institute have introduced FutureSim, a new benchmark designed to evaluate AI agents' ability to predict real-world events using only historical web data. This method prevents agents from accessing future information, simulating a more realistic forecasting scenario. Early tests using models like GPT-5.5 within the Codex harness showed strong performance on some markets, such as the Super Bowl, but struggled with others like UK elections and the Grammys, indicating narrow capabilities. AI

IMPACT Tests AI agents' ability to forecast events using historical data, highlighting narrow capabilities beyond trivia.

RANK_REASON The cluster describes the release of a new academic benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]

Read on dev.to — Claude Code tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

FutureSim benchmark tests AI forecasting with historical data

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes the release of a new academic benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
120 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — Claude Code tag TIER_1 English(EN) · Simon Paxton ·

    FutureSim Exposes Polymarket AI's Narrow Wins and Failures

    <p>Max Planck Institute researchers recently released FutureSim, a benchmark for polymarket ai-style forecasting that tests whether agents can predict real-world events from a frozen slice of past web history rather than the live internet.</p> <p>According to the FutureSim projec…