PulseAugur
中
实时 12:16:55
English(EN) LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization

新方法通过推测解码加速大语言模型推理 · 跟踪 7 个来源

研究人员正在开发新的方法,通过推测解码来加速大语言模型(LLM)的推理速度。DARTree 和 SPADE 是两种此类方法,其中 DARTree 专注于基于树的推测解码,以提高接受长度和加速效果,而 SPADE 则将推测解码集成到边缘和云设备中,以降低成本和延迟。其他相关工作包括 MemSpec,它针对内存受限的边缘设备优化自适应推测解码,以及 Goose,它使用各向异性推测树来提高效率。这些进展旨在使大语言模型的部署更加实用和经济高效。 AI

影响 加速大语言模型推理,可能降低人工智能应用的部署成本和延迟。

排序理由 多篇研究论文介绍了大语言模型推测解码的新方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

新方法通过推测解码加速大语言模型推理 · 跟踪 7 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了大语言模型推测解码的新方法。
Source corroboration
9 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
55 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [9]

  1. arXiv cs.CL TIER_1 English(EN) · Haoxuan Luo, Jameson Sandler, Ferdinando Fioretto ·

    从逐点置信度到前缀调度:投机解码中的验证器跳过

    arXiv:2608.14787v1 Announce Type: cross Abstract: Speculative decoding is a leading technique to reduce the cost of autoregressive generation by using a small drafter to propose several tokens, which are then verified in parallel by a larger target model. Speculative diffusion de…

  2. arXiv cs.AI TIER_1 English(EN) · Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal ·

    SPADE:用于精确低成本分布式边缘云推理的推测解码

    arXiv:2608.13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circu…

  3. arXiv cs.LG TIER_1 English(EN) · Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen ·

    DARTree:具有自回归草稿树的推测性扩散解码

    arXiv:2608.13524v1 Announce Type: new Abstract: Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    DARTree:具有自回归草稿树的推测性扩散解码

    Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal …

  5. arXiv cs.LG TIER_1 English(EN) · Pranav Subbaraman, Fang Sun, Jinxi Yu, Yue Yao, Huacong Tang, Xiao Luo, Yizhou Sun ·

    使用推测解码加速时间序列基础模型

    arXiv:2511.18191v2 Announce Type: replace Abstract: Time series forecasting drives operational decisions under tight latency budgets, and autoregressive time series foundation models (TSFMs) increasingly deliver the most accurate forecasts. That accuracy is paid for at inference,…

  6. arXiv cs.AI TIER_1 English(EN) · Eunjeong Kim, Yeong Jun Jeon, Myeonggyun Han ·

    MemSpec:内存感知运行时,用于边缘设备投机解码中的自适应草稿调度

    arXiv:2608.10362v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps. Its effectiveness depends heavil…

  7. arXiv cs.AI TIER_1 English(EN) · Zexun Lin, Yuan Feng, Junlin Lv, Kevin S. Zhou, Xike Xie ·

    LibraSpec:基于边际增益驱动优化的动态扩散式推测解码

    arXiv:2608.08721v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynam…

  8. arXiv cs.AI TIER_1 English(EN) · Tao Jin, Phuong Minh Nguyen, Naoya Inoue ·

    Goose:用于无训练推测解码的各向异性推测树

    arXiv:2604.02047v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forward pass. Candidates are organized as a tree: deeper trees accept more tokens per ste…

  9. Towards AI TIER_1 English(EN) · Armin Norouzi, Ph.D ·

    使用 Python 中的草稿模型实现推测解码

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/implement-speculative-decoding-with-a-draft-model-in-python-bd2e7e6b483b?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1169/1*XuSjVweq3CXXV0jG9az1Rw.png" …