Two new research papers explore enhancements to Mixture-of-Experts (MoE) large language models. The first paper introduces token-error supervision to guide MoE routing, improving accuracy on benchmarks like Granite and ARC-Challenge by aligning routing affinities with actual token errors. The second paper proposes Enhanced Mixture-of-Experts (E-MoE) for non-factorized diffusion language models, using MoE routing to capture cross-position correlations in a discrete latent space without increasing active parameters, leading to better few-step generation quality. AI
IMPACT These papers introduce novel techniques for improving the efficiency and performance of Mixture-of-Experts models, potentially leading to more capable and faster LLMs.
RANK_REASON Two academic papers published on arXiv detailing novel methods for Mixture-of-Experts large language models.
- ARC challenge
- arXiv
- cross entropy
- Granite
- Hugging Face
- Itakura--Saito divergence
- large-language models
- LM1B
- Masked Diffusion Models
- mixture of experts
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →