A developer has created a custom C++ engine and a new 4-bit quantization format, H128/Q4-G32-DOT4, for the Qwen 3.5 0.8B model. This new format results in a smaller model size of 425 MB, which is 71 MB less than Unsloth's mixed-precision Q4_0. Benchmarks on a Ryzen 9 9955HX3D CPU show significant performance improvements, with the custom engine achieving up to 2.9x faster prefill and 1.7x higher batch-16 throughput compared to other engines. AI
IMPACT Enables more efficient local deployment of smaller language models on consumer hardware.
RANK_REASON Custom engine and quantization format for an existing model.
- 0.8B
- central processing unit
- H128/Q4-G32-DOT4
- ik_llama.cpp
- llama.cpp
- Qwen 3.5
- Ryzen 9 9955HX3D
- Unsloth
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →