deepseek-moe-16b-base
PulseAugur coverage of deepseek-moe-16b-base — every cluster mentioning deepseek-moe-16b-base across labs, papers, and developer communities, ranked by signal.
-
New method prunes MoE language models using generic text corpora
Researchers have developed a new method called Generic TB-Coverage for pruning sparsely activated Mixture-of-Experts (MoE) language models. This technique addresses the challenge of removing redundant experts without re…
-
New MoE Pruning Method Uses Generic Data to Preserve Expert Utility
Researchers have developed a new method called Generic TB-Coverage for pruning sparsely activated Mixture-of-Experts (MoE) language models. This approach uses generic text corpora like WikiText2 and C4 for calibration, …
-
ConMoE framework compresses MoE models without retraining
Researchers have developed ConMoE, a novel framework for compressing Mixture-of-Experts (MoE) language models without requiring retraining. This method consolidates the expert pool by reassigning original expert referen…