DFlash 2
PulseAugur coverage of DFlash 2 — every cluster mentioning DFlash 2 across labs, papers, and developer communities, ranked by signal.
- 2026-08-18 research_milestone DFlash 2, an updated quantization technique, has been released and is available for Qwen 3.8-27B and Muse Glimmer models. source
-
ik_llama.cpp adds DFlash 2 speculative decoding and new model support
A pull request has been submitted to the ik_llama.cpp project, introducing DFlash 2 speculative decoding. This update also includes support for IQ4_KS and IQ4_KT quantization formats on RDNA3 GPUs via HIP, Vulkan, and i…
-
DFlash 2 boosts Qwen 3.8 27B speed by 2.26x in llama.cpp benchmarks
A user has benchmarked the new DFlash 2 speculative decoding method within llama.cpp, using the Qwen 3.8 27B model. The results show a 2.26x speed increase on real-world coding prompts without additional methods, and up…
-
Inco AI releases DFlash 2 for faster LLM inference
Inco AI has released DFlash 2, an advancement in speculative decoding for large language models. This new version improves output by over 20% per verification pass with minimal latency increase, building on the original…
-
DFlash 2 quantization technique released for Qwen 3.8-27B and Muse Glimmer
DFlash 2, an updated version of the DFlash quantization technique, has been released. This new version is available for the Qwen 3.8-27B and Muse Glimmer large language models. GGUF quantizations for DFlash 2 have alrea…