PulseAugur
EN
LIVE 00:04:39

RAZOR method prunes LLM experts without sacrificing reasoning ability

Researchers have developed RAZOR, a novel method for pruning experts in Mixture-of-Experts (MoE) models without significantly degrading their reasoning capabilities. Unlike previous methods that focus on expert frequency or isolated contribution, RAZOR assesses functional replaceability by measuring the damage caused by expert deletion. This training-free approach uses consensus residuals to calculate the exact output change from removing an expert, enabling pruning with forward passes alone. RAZOR has demonstrated superior performance across multiple LLMs and expert removal budgets, outperforming existing methods on reasoning-centered tasks. AI

IMPACT This method could lead to more efficient LLMs by reducing computational requirements without compromising performance on complex reasoning tasks.

RANK_REASON The cluster contains a research paper detailing a new method for pruning LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

RAZOR method prunes LLM experts without sacrificing reasoning ability

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Mingyang Song, Mao Zheng ·

    RAZOR: Pruning Replaceable Experts in LLMs

    arXiv:2609.30465v2 Announce Type: new Abstract: Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Commo…