PulseAugur
中
实时 09:43:26
English(EN) Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

新的训练后方法以更少的计算量提升大语言模型吞吐量

研究人员开发了一种新的语言模型多令牌预测(MTP)训练后方法,与传统的联合训练相比,显著降低了计算成本。该技术允许一个固定的推理模型在数学和编码等基准测试上,使用少得多的训练令牌,实现相当或甚至更好的吞吐量速度。该方法包括对草稿令牌验证规则进行链式感知松弛,以允许从骨干分布中有限的漂移,以及一个自适应控制器,在推理过程中动态调整MTP头的数量,恢复速度损失。 AI

影响 这项研究可能通过降低实现高生成吞吐量所需的计算要求,从而实现更高效的大语言模型部署。

排序理由 关于语言模型多令牌预测新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的训练后方法以更少的计算量提升大语言模型吞吐量

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
关于语言模型多令牌预测新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song ·

    匹配分布,而非计算:训练后多令牌预测头

    arXiv:2610.00888v1 Announce Type: cross Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distri…