PulseAugur
EN
LIVE 20:46:38

Liquid AI releases draft models for 3.18x faster decoding

Liquid AI has introduced DSpark draft models for its LFM2.5 family, designed to accelerate decoding speeds by up to 3.18x without altering the final model outputs. These models utilize a speculative decoding approach, where a smaller "drafter" model proposes tokens that a larger "target" model then verifies in a single pass. This method offers significant speedups, particularly for applications involving reasoning and tool calls, and is supported by tools like llama.cpp and SGLang. The models are available under the LFM Open License v1.0, which permits free commercial use for entities with annual revenue under $10 million. AI

IMPACT Accelerates inference for LLM applications, particularly agentic workflows, by reducing latency without compromising output quality.

RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on MarkTechPost →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Liquid AI releases draft models for 3.18x faster decoding

COVERAGE [1]

  1. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs

    <p>Three ~300M drafters bring speculative decoding to LFM2.5, delivering up to 3.18x faster decoding with identical greedy output.</p> <p>The post <a href="https://www.marktechpost.com/2026/08/20/liquid-ai-releases-lfm2-5-dspark-draft-models-that-deliver-up-to-3-18x-faster-decodi…