A new release of Qwen3.8-Flash-Next models is available, featuring advanced quantization techniques like GSQ and RCO. These models offer significant reductions in size while maintaining performance, with one version achieving 93.26% of the base model's performance at 3.50 bits per parameter. A specialized 'Coder' build further optimizes by removing half of the model's experts, resulting in a 29.6 GB resident working set that can run on a single 32 GB accelerator, while still retaining over 90% of its coding benchmark capabilities. AI
IMPACT Offers highly compressed models for efficient deployment on consumer hardware, enabling broader access to advanced AI capabilities.
RANK_REASON Release of quantized and pruned open-source models with technical details on quantization methods. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →