PulseAugur
EN
LIVE 05:04:21

New benchmark UnpredictaBench tests LLM distributional randomness

Researchers have introduced UnpredictaBench, a new benchmark designed to evaluate how well large language models (LLMs) can capture true underlying probability distributions. The benchmark addresses the issue of LLMs collapsing to a single plausible answer, which is problematic for simulations requiring calibrated samples. UnpredictaBench includes 448 problems across various distributions and uses the KS@N metric to quantify model performance, revealing a wide range of distributional capabilities among tested models. AI

IMPACT Highlights a critical gap in LLM capabilities for simulation and complex system modeling, suggesting further research is needed for true distributional sampling.

RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

New benchmark UnpredictaBench tests LLM distributional randomness

COVERAGE [3]

  1. arXiv cs.CL TIER_1 English(EN) · Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Liang Luo, Ellie Dingqiao Wen, Lele Wang, Giuseppe Carenini, Peter West ·

    UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

    arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are increasingly used as substitutes for other entities (e.g., for humans in econom…

  2. arXiv cs.CL TIER_1 English(EN) · Peter West ·

    UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

    We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are increasingly used as substitutes for other entities (e.g., for humans in economic simulations), the tendency of many models to …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

    UnpredictaBench evaluates large language models' capacity to sample from target distributions, revealing significant gaps in their ability to simulate unpredictable systems despite recent advances in output diversity.