Qwen1.5-MoE-A2.7B
PulseAugur coverage of Qwen1.5-MoE-A2.7B — every cluster mentioning Qwen1.5-MoE-A2.7B across labs, papers, and developer communities, ranked by signal.
-
New method prunes MoE language models using generic text corpora
Researchers have developed a new method called Generic TB-Coverage for pruning sparsely activated Mixture-of-Experts (MoE) language models. This technique addresses the challenge of removing redundant experts without re…
-
New MoE Pruning Method Uses Generic Data to Preserve Expert Utility
Researchers have developed a new method called Generic TB-Coverage for pruning sparsely activated Mixture-of-Experts (MoE) language models. This approach uses generic text corpora like WikiText2 and C4 for calibration, …
-
AI research questions expert importance metrics in MoE models
A new research paper investigates the effectiveness of interpretability methods in Mixture-of-Experts (MoE) models. The study found that common metrics used to predict which experts can be removed without impacting perf…