PulseAugur
实时 10:28:42
English(EN) Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

研究发现:长上下文训练可能损害大型语言模型的知识

一篇新研究论文提出了“信息丰富悖论”,挑战了大型语言模型中更长上下文窗口总是能提高性能的假设。研究表明,训练过程中过多的相关信息会降低模型进行参数化编码知识的动力,导致过度依赖上下文。这种现象在语言建模、自然语言理解和闭卷问答任务中,超过某个中间最优值后,会降低性能。研究表明,使用信息丰富的上下文进行训练会将梯度压力从与参数化知识相关的馈送网络转移到注意力模块,从而在推理过程中增加上下文依赖性。 AI

影响 挑战了长上下文窗口普遍提高大型语言模型性能的假设,暗示了参数化知识与上下文依赖性之间可能存在的权衡。

排序理由 一篇介绍大型语言模型训练新悖论和假设的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:长上下文训练可能损害大型语言模型的知识

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Arda Uzunoglu, Benjamin van Durme, Daniel Khashabi ·

    信息爆炸悖论:长上下文训练损害参数化知识

    arXiv:2608.12218v1 Announce Type: cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will …