Researchers have introduced SocietyBench, a novel benchmark designed to evaluate how well large language models (LLMs) can forecast counterfactual social-world evolution. Unlike existing benchmarks that focus on task completion, SocietyBench assesses an LLM's understanding of social dynamics by presenting it with timelines of factual events and public opinion, then asking forecasting questions. To prevent models from relying on pre-training memory, named entities and dates are altered, creating a structurally similar but counterfactual scenario. The benchmark includes both Chinese and English editions, with the strongest LLM achieving only 75.0 out of 100 on two orthogonal axes: probability calibration and temporal accuracy, indicating a significant gap in current LLM capabilities. AI
IMPACT This benchmark could drive the development of LLMs with improved reasoning and forecasting abilities in complex social scenarios.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →