PulseAugur
EN
LIVE 03:31:02

Inception Labs launches Mercury 2.5 diffusion LLM at 1,107 tokens/sec · 1 source tracked

Inception Labs has unveiled Mercury 2.5, a diffusion language model that achieves 1,107 tokens per second on NVIDIA GPUs. This represents a significant speed increase over traditional autoregressive models like GPT-3, which generate text token by token. Mercury 2.5 utilizes a diffusion generation process, similar to image generation models like Stable Diffusion, allowing for parallel refinement of the entire output rather than sequential token prediction. The company claims this approach offers a 40% intelligence gain over its predecessor, Mercury 2, and provides a tunable AI

IMPACT This diffusion-based approach could significantly accelerate LLM inference speeds, potentially enabling new real-time applications.

RANK_REASON The item describes a new model release from a lab (Inception Labs) with a specific name and performance metric. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Inception Labs launches Mercury 2.5 diffusion LLM at 1,107 tokens/sec · 1 source tracked

How we ranked this

Signal score
17 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
The item describes a new model release from a lab (Inception Labs) with a specific name and performance metric. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Daniel Sam Pete Thiyagu ·

    1,107 Tokens Per Second: The LLM That Doesn't Type

    <p>On September 8, 2026, Inception Labs announced <strong>Mercury 2.5</strong> — which the company describes as the largest diffusion language model ever trained. The headline number: <strong>1,107 tokens per second</strong> on widely available NVIDIA GPUs, at quality the company…