A new Mixture of Experts (MoE) model, Whittle MoE 27B, has been developed by carving experts from the Qwen3.8-27B model and retraining only the routers. This approach aims to reduce the model's tendency to loop and truncate responses, with reported improvements in conversational ability and structured output generation. The model requires 24GB of VRAM for quantized operation and is available for download via Hugging Face, though further development is dependent on donations. AI
IMPACT Demonstrates a method for improving existing models via router retraining, potentially offering a path to more efficient fine-tuning.
RANK_REASON Model release from a non-frontier lab with detailed technical description. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →