A new version of the GLM-5.2 model, named "colibri int4 with int8 mtp", has been released on Hugging Face. This iteration is based on the original GLM-5.2 model and features int8 MTP heads designed to significantly boost inference speed through speculative decoding. The model requires the colibri engine for operation and is not compatible with standard formats like GGUF or AWQ. It is distributed under the MIT license. AI
IMPACT Offers a potential speed boost for inference, requiring specific engine compatibility.
RANK_REASON Release of a specific model variant with performance improvements, not from a frontier lab. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
- colibri
- GLM-5.2
- jlnsrk/GLM-5.2-colibri-int4
- mateogrgic/GLM-5.2-colibri-int4-with-int8-mtp
- zai-org/GLM-5.2-FP8
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →