PulseAugur
实时 05:42:30

新基准 UnpredictaBench 测试大型语言模型分布随机性

研究人员推出了 UnpredictaBench,这是一个旨在评估大型语言模型(LLM)捕捉真实底层概率分布能力的新基准。该基准解决了 LLM 倾向于生成单一合理答案的问题,这对于需要校准样本的模拟来说是有问题的。UnpredictaBench 包含跨越各种分布的 448 个问题,并使用 KS@N 指标来量化模型性能,揭示了不同模型之间分布能力的巨大差异。 AI

影响 突显了 LLM 在模拟和复杂系统建模能力方面的一个关键差距,表明需要进一步研究以实现真正的分布采样。

排序理由 该集群包含一篇介绍用于评估 LLM 的新基准的研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新基准 UnpredictaBench 测试大型语言模型分布随机性

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Liang Luo, Ellie Dingqiao Wen, Lele Wang, Giuseppe Carenini, Peter West ·

    UnpredictaBench:用于评估大型语言模型中分布随机性的基准测试

    arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are increasingly used as substitutes for other entities (e.g., for humans in econom…

  2. arXiv cs.CL TIER_1 English(EN) · Peter West ·

    UnpredictaBench:用于评估大型语言模型中分布随机性的基准测试

    We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are increasingly used as substitutes for other entities (e.g., for humans in economic simulations), the tendency of many models to …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    UnpredictaBench:用于评估大型语言模型中分布随机性的基准测试

    UnpredictaBench evaluates large language models' capacity to sample from target distributions, revealing significant gaps in their ability to simulate unpredictable systems despite recent advances in output diversity.