PulseAugur
实时 02:23:20
English(EN) Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

新基准显示大型语言模型在复杂科学推理方面存在困难

开发了两个新基准 PHDSDABench,用于评估大型语言模型 (LLM) 在科学发现和数据分析方面的能力。PHD 侧重于模型从不确凿证据中自主构建假设空间的能力,而 SDABench 则评估了包括探索、推理和跨不同科学领域的机械推理在内的六项特定科学能力。对 15 个大型语言模型的初步评估表明,虽然模型在描述性任务方面表现出色,但在更复杂的推理、假设选择和机械解释方面存在困难,这表明它们在科学发现潜力方面存在显著差距。 AI

影响 突出了当前大型语言模型在复杂科学推理方面的局限性,指出了未来模型开发的方向。

排序理由 该集群描述了用于评估大型语言模型在科学发现和数据分析方面的新学术基准。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新基准显示大型语言模型在复杂科学推理方面存在困难

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Tianyun Zhong, Wangyi Jiang, Wei Wang, Xuanang Chen, Yaojie Lu, Shiwei Ye, Yuzhen Shi, Boyu Yang, Jinghang Wang, Han Li, Weiqi Zhai, Bing Zhao, Hu Wei, Haiyang Yu, Yongbin Li, Hongyu Lin, Le Sun, Xianpei Han ·

    行动之前:在预期假设发现方面对大型语言模型进行基准测试

    arXiv:2607.15766v1 Announce Type: new Abstract: Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured. We introduce Prospective Hypothesis Discovery (PHD)…

  2. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    SDABench:用于评估 LLM 在科学发现中表现的新基准

    <h2> What Changed </h2> <p>Traditional benchmarks for evaluating Large Language Models (LLMs) in scientific data analysis have primarily focused on code execution or workflow completion. This approach often fails to account for the distinct types of scientific claims that analysi…