PulseAugur
中
实时 04:21:04
English(EN) 1,107 Tokens Per Second: The LLM That Doesn't Type

Inception Labs 发布 Mercury 2.5 扩散式 LLM,速度达每秒 1,107 个 token · 跟踪 1 个来源

Inception Labs 推出了 Mercury 2.5,这是一款在 NVIDIA GPU 上实现每秒 1,107 个 token 的扩散式语言模型。与像 GPT-3 这样逐个 token 生成文本的传统自回归模型相比,这代表了显著的速度提升。Mercury 2.5 采用扩散生成过程,类似于 Stable Diffusion 等图像生成模型,允许对整个输出进行并行优化,而不是顺序预测 token。该公司声称,这种方法比其前身 Mercury 2 提供了 40% 的智能提升,并提供了一个可调的 AI

影响 这种基于扩散的方法可以显著加快 LLM 的推理速度,可能催生新的实时应用。

排序理由 该条目描述了一个实验室(Inception Labs)发布的新模型,并附有具体名称和性能指标。[lever_c_从 frontier_release 降级:ic=1 ai=1.0]

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Inception Labs 发布 Mercury 2.5 扩散式 LLM,速度达每秒 1,107 个 token · 跟踪 1 个来源

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
该条目描述了一个实验室(Inception Labs)发布的新模型,并附有具体名称和性能指标。[lever_c_从 frontier_release 降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Daniel Sam Pete Thiyagu ·

    每秒1107个Token:这个LLM不会打字

    <p>On September 8, 2026, Inception Labs announced <strong>Mercury 2.5</strong> — which the company describes as the largest diffusion language model ever trained. The headline number: <strong>1,107 tokens per second</strong> on widely available NVIDIA GPUs, at quality the company…