Researchers have developed RAZOR, a novel method for pruning experts in Mixture-of-Experts (MoE) models without significantly degrading their reasoning capabilities. Unlike previous methods that focus on expert frequency or isolated contribution, RAZOR assesses functional replaceability by measuring the damage caused by expert deletion. This training-free approach uses consensus residuals to calculate the exact output change from removing an expert, enabling pruning with forward passes alone. RAZOR has demonstrated superior performance across multiple LLMs and expert removal budgets, outperforming existing methods on reasoning-centered tasks. AI
IMPACT This method could lead to more efficient LLMs by reducing computational requirements without compromising performance on complex reasoning tasks.
RANK_REASON The cluster contains a research paper detailing a new method for pruning LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- DeepSeek-V4-Flash-0731
- GLM 4.7 Flash
- Mingyang Song
- Mixture-of-Experts (MoE) models
- Qwen3.6 35B-A3B
- RAZOR
- REAP
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →