Researchers from the University of California, San Diego, have developed DFlash, a novel speculative decoding technique that significantly accelerates AI inference. Unlike traditional methods that generate tokens one by one, DFlash uses a block diffusion model to propose entire blocks of tokens in a single pass. A larger target model then verifies these blocks in parallel, leading to substantial speedups. This approach has demonstrated up to 15x higher throughput on NVIDIA Blackwell GPUs for models like GPT-OSS 120B, making it particularly beneficial for long-context reasoning and coding tasks. AI
IMPACT Accelerates AI inference speed, potentially reducing costs and improving user experience for generative AI applications.
RANK_REASON Research paper detailing a new technique for speculative decoding in LLMs.
Read on Mastodon — fosstodon.org →
- DiffuSpec
- GPT-OSS 120B
- NVIDIA
- NVIDIA Blackwell
- Qwen3-coder
- SpecDiff-2
- University of California, San Diego
- Speculative decoding
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →