PulseAugur
中
实时 23:33:46
English(EN) Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding

新的H2S模型通过压缩证据在长上下文理解方面表现出色

研究人员推出了一种用于大型语言模型长上下文理解的新范式——Highlight-Then-Summarize (H2S)。该方法首先从长文档中识别相关证据,然后将其压缩成摘要,最后生成最终答案。研究人员开发了包含超过6600个示例的H2S-Dataset和一种强化学习方法H2S-RL,以训练这种“先压缩后推理”的行为。在H2S-Bench上的评估表明,H2S-14B模型在证据选择、摘要质量和有限输出预算内的最终答案准确性方面,显著优于其他开源模型,并取得了优异的成绩。 AI

影响 该方法可以通过关注相关证据来提高LLM处理长文档的效率和准确性。

排序理由 这是一篇详细介绍LLM长上下文理解新方法和新数据集的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的H2S模型通过压缩证据在长上下文理解方面表现出色

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc), Gaofeng Lu (Baidu Inc), Zhenyu Jiao (Baidu Inc), Xing Yuan (Baidu Inc), Dai Dai (Baidu Inc), Tong Mo (… ·

    高亮后总结:学习压缩证据以实现长上下文理解

    arXiv:2609.31382v1 Announce Type: cross Abstract: Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content…