Inception Labs has released Mercury 2.5, a new diffusion LLM that offers a 40% increase in intelligence over its predecessor, Mercury 2. This model boasts impressive speed at 1,107 tokens per second on NVIDIA GPUs and supports a 260K token context window. Mercury 2.5 is positioned as a cost-effective alternative, comparable to models like GPT-5.6 Luna (Low) and Gemini 3.5 Flash-Lite, with introductory pricing significantly reduced. The release also includes previews of Mercury Voice and Mercury Router, aimed at further optimizing latency-sensitive applications and intelligent model routing. AI
IMPACT Sets a new benchmark for diffusion LLMs in terms of intelligence, speed, and cost-effectiveness, potentially impacting search, voice, and coding applications.
RANK_REASON New model release from a frontier lab (Inception Labs) with performance metrics and comparisons to other frontier models. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Hacker News — AI stories ≥50 points →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →