Two new research papers propose novel methods for pruning Mixture-of-Experts (MoE) language models to reduce memory usage without sacrificing performance. The first paper introduces AIMER, a calibration-free criterion that ranks experts based on the concentration of their weights, outperforming existing methods on various benchmarks and model sizes. The second paper offers a unified formulation for one-shot MoE expert pruning, leading to a selection principle for task-agnostic versus task-specific pruning and introducing two new criteria, MAN and MSAN, which show strong performance across multiple models and tasks. AI
IMPACT These methods could significantly reduce the memory footprint of large MoE models, making them more accessible and efficient for deployment.
RANK_REASON Two academic papers published on arXiv proposing new methods for MoE expert pruning.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Hugging Face
- IArxiv
- Mean Activation Norm
- Mean Squared Activation Norm
- Mixture-of-Experts
- AIMER
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →