PulseAugur
实时 05:59:23
English(EN) Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

新基准测试大型语言模型从书目中推断研究思路的能力

研究人员推出 Reconstruction,一个旨在评估语言模型仅从预发布书目中推断研究思路能力的新型基准。该基准采用严格的防泄露协议,包括时间引用截止和匿名参考文献ID,以确保评估的完整性。虽然前沿模型在六个科学领域和643篇论文上的匹配率仅为3-15%,但结合了跨模型审查和锦标赛结构的多个智能体管道显著提高了性能,匹配率达到了23-42%。 AI

影响 该基准有望推动大型语言模型更复杂的推理和信息提取能力的发展。

排序理由 该集群描述了一篇在arXiv上发表的新基准和研究论文,详细介绍了一种评估语言模型的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.MA (Multiagent) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准测试大型语言模型从书目中推断研究思路的能力

报道来源 [1]

  1. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ritankar Das ·

    Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

    Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks mo…