Liquid AI has introduced DSpark draft models for its LFM2.5 family, designed to accelerate decoding speeds by up to 3.18x without altering the final model outputs. These models utilize a speculative decoding approach, where a smaller "drafter" model proposes tokens that a larger "target" model then verifies in a single pass. This method offers significant speedups, particularly for applications involving reasoning and tool calls, and is supported by tools like llama.cpp and SGLang. The models are available under the LFM Open License v1.0, which permits free commercial use for entities with annual revenue under $10 million. AI
IMPACT Accelerates inference for LLM applications, particularly agentic workflows, by reducing latency without compromising output quality.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- DSpark
- GGUF
- Hugging Face
- LFM2.5
- LFM2.5-1.2B-Instruct
- LFM2.5-2.6B
- LFM2.5-8B-A1B
- LFM Open License v1.0
- Liquid AI
- llama.cpp
- M4 Max MacBook Pro
- NVIDIA H100
- safetensors
- SGLang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →