PulseAugur
EN
LIVE 00:04:36

Qwen3.8-Flash-Next models released with expert pruning and advanced quantization

A new release of Qwen3.8-Flash-Next models is available, featuring advanced quantization techniques like GSQ and RCO. These models offer significant reductions in size while maintaining performance, with one version achieving 93.26% of the base model's performance at 3.50 bits per parameter. A specialized 'Coder' build further optimizes by removing half of the model's experts, resulting in a 29.6 GB resident working set that can run on a single 32 GB accelerator, while still retaining over 90% of its coding benchmark capabilities. AI

IMPACT Offers highly compressed models for efficient deployment on consumer hardware, enabling broader access to advanced AI capabilities.

RANK_REASON Release of quantized and pruned open-source models with technical details on quantization methods. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8-Flash-Next models released with expert pruning and advanced quantization

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Loginhe ·

    [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wt4s88/release_gsqrco_ggufs_for_qwen38flashnext_plus_a/"> <img alt="[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw" src="https://preview.redd.it/penwbxul8fsh…