PulseAugur
EN
LIVE 08:53:54

New SocietyBench benchmark tests LLMs on forecasting counterfactual social events

Researchers have introduced SocietyBench, a novel benchmark designed to evaluate how well large language models (LLMs) can forecast counterfactual social-world evolution. Unlike existing benchmarks that focus on task completion, SocietyBench assesses an LLM's understanding of social dynamics by presenting it with timelines of factual events and public opinion, then asking forecasting questions. To prevent models from relying on pre-training memory, named entities and dates are altered, creating a structurally similar but counterfactual scenario. The benchmark includes both Chinese and English editions, with the strongest LLM achieving only 75.0 out of 100 on two orthogonal axes: probability calibration and temporal accuracy, indicating a significant gap in current LLM capabilities. AI

IMPACT This benchmark could drive the development of LLMs with improved reasoning and forecasting abilities in complex social scenarios.

RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New SocietyBench benchmark tests LLMs on forecasting counterfactual social events

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi ·

    SocietyBench: Forecasting Counterfactual Social-World Evolution

    arXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model u…