Inception Labs has unveiled Mercury 2.5, a diffusion language model that achieves 1,107 tokens per second on NVIDIA GPUs. This represents a significant speed increase over traditional autoregressive models like GPT-3, which generate text token by token. Mercury 2.5 utilizes a diffusion generation process, similar to image generation models like Stable Diffusion, allowing for parallel refinement of the entire output rather than sequential token prediction. The company claims this approach offers a 40% intelligence gain over its predecessor, Mercury 2, and provides a tunable AI
IMPACT This diffusion-based approach could significantly accelerate LLM inference speeds, potentially enabling new real-time applications.
RANK_REASON The item describes a new model release from a lab (Inception Labs) with a specific name and performance metric. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Artificial Analysis
- Claude Haiku-4-5
- Gemini 3.8 Flash
- GPT-3
- Inception Labs
- Mercury 2
- Mercury 2.5
- Nvidia
- Stable Diffusion
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →