PulseAugur
实时 09:10:42
English(EN) Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

研究发现:看似微小的架构选择严重影响大语言模型长上下文扩展

一篇新发表在arXiv上的研究论文详细介绍了Transformer模型中看似微小的架构选择如何显著影响其扩展上下文长度的能力。研究发现,结合三个或更多特定的架构决策(存在于Olmo、Llama和Qwen等模型中)可以将长上下文性能降低高达47%。这些差异在标准的短上下文验证中无法检测到,但在预训练早期应用上下文扩展时会显现出来。研究人员发布了一套名为OlmPool的26个可比的7B模型,以促进进一步研究,其中一些架构在长上下文扩展性方面优于Llama 3。 AI

影响 强调了长上下文性能的关键架构因素,可能指导未来的模型开发。

排序理由 研究论文,详细介绍了关于大语言模型架构和长上下文扩展的发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:看似微小的架构选择严重影响大语言模型长上下文扩展

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Amanda Bertsch, Luca Soldaini, Matthew R. Gormley, Graham Neubig, Hannaneh Hajishirzi, Kyle Lo, Dirk Groeneveld ·

    基础的裂痕:看似微小的架构选择影响长上下文扩展

    arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set…