PulseAugur
中
实时 09:38:08
English(EN) SEED: Self-Speculative Decoding via Implicit Encoder-Decoder

新的推测解码方法提高了LLM推理效率 · 跟踪4个来源

研究人员开发了几种新方法来提高大型语言模型中推测解码的效率。DSpine在模型的骨干网络中引入因果条件注入,以增强令牌之间的信息流,在Qwen3模型上实现了显著的加速和更长的接受长度。LongSpark提出了一种固定成本的并行草稿生成器,使解码成本独立于前缀长度,从而在长上下文任务上提高效率。DScale通过使用自适应验证、路径感知瓦片和动态长度分配来扩展块扩散推测解码,从而在吞吐量方面比现有方法有显著提升。SEED将Transformer重新解释为隐式编码器-解码器,通过重用计算出的表示来廉价地实现高质量的草稿生成,从而提高速度和生成质量。 AI

影响 这些在推测解码方面的进展可以显著降低大型语言模型的推理成本和延迟,从而实现更广泛的应用和更高效的实时应用。

排序理由 arXiv上发表了多篇研究论文,详细介绍了大型语言模型中推测解码的新颖方法。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新的推测解码方法提高了LLM推理效率 · 跟踪4个来源

本文如何被排名

Signal score
25 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv上发表了多篇研究论文,详细介绍了大型语言模型中推测解码的新颖方法。
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [4]

  1. arXiv cs.AI TIER_1 English(EN) · Haohui Zhang, Keyu Chen, Haocheng Sun, Weibo Gu, Ruizhi Qiao, Xing Sun, Bo Jiang ·

    并行草稿,深度条件:用于推测解码的相邻因果注入

    arXiv:2609.36173v1 Announce Type: cross Abstract: Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional d…

  2. arXiv cs.AI TIER_1 English(EN) · Hao-Yuan He, Peng-Fei Liu, Si Shen, Ming Li ·

    LongSpark:具有固定成本并行草稿器的有效推测解码

    arXiv:2609.37029v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding…

  3. arXiv cs.AI TIER_1 English(EN) · Rongjian Chen, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu ·

    DScale:通过自适应验证扩展块扩散投机解码

    arXiv:2609.37532v1 Announce Type: cross Abstract: Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes …

  4. arXiv cs.CL TIER_1 English(EN) · Hankun Lin, Patrick Pynadath, Ruqi Zhang ·

    SEED:通过隐式编码器-解码器进行自我推测解码

    arXiv:2609.36590v1 Announce Type: new Abstract: Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts chea…