Liquid AI has released DSpark draft models for its LFM2.5 series, which enhance decoding speed by up to 3.18x without altering output quality. These models utilize speculative decoding, where a smaller draft model proposes token candidates that a larger target model verifies. This approach significantly speeds up inference, particularly for agentic applications where latency is critical. The models are available for self-hosting and are compatible with tools like llama.cpp and SGLang. AI
IMPACT Accelerates inference for smaller models, potentially enabling more complex AI applications on edge devices.
RANK_REASON Liquid AI is a frontier lab releasing new draft models with a novel speculative decoding technique. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- DSpark
- GGUF
- Hugging Face
- LFM2.5
- LFM2.5-1.2B-Instruct
- LFM2.5-2.6B
- LFM2.5-8B-A1B
- LFM Open License v1.0
- Liquid AI
- llama.cpp
- M4 Max MacBook Pro
- NVIDIA H100
- safetensors
- SGLang
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →