PulseAugur
实时 15:47:02
English(EN) Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models

Dolphin模型为长上下文AI提供10倍能效

研究人员推出了一种新颖的解码器-解码器架构Dolphin,旨在实现设备端语言模型中长上下文的高效处理。该方法使用一个较小的解码器将广泛的上下文提炼成内存嵌入,从而缩短主解码器的输入长度。通过将长文本视为一种独特的模态,类似于图像嵌入,Dolphin在不影响响应质量的情况下,实现了十倍的能效提升和五倍的延迟降低。该模型已在Hugging Face上公开可用,旨在在资源受限的环境中实现更复杂的AI能力。 AI

影响 通过提高长上下文处理的能效和降低延迟,在边缘设备上实现更复杂的AI能力。

排序理由 该集群包含一篇详细介绍新模型架构的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Dolphin模型为长上下文AI提供10倍能效

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Wei Chen, Zhiyuan Li, Shuo Xin, Yihao Wang ·

    Squid:长上下文作为一种新的模态,用于高能效的设备端语言模型

    arXiv:2408.15518v3 Announce Type: replace Abstract: This paper presents Dolphin, a novel decoder-decoder architecture for energy-efficient processing of long contexts in language models. Our approach addresses the significant energy consumption and latency challenges inherent in …