Two new research papers explore the complexities of Mixture-of-Experts (MoE) models, particularly concerning calibration and discontinuities. The first paper investigates how expert-level calibration impacts MoE performance under distribution shifts, proposing an adversarial reweighting method to improve accuracy and calibration for soft-routed models. The second paper provides a rigorous geometric and stochastic analysis of discontinuities in Sparse Mixture-of-Experts (SMoE) architectures, identifying that lower-order discontinuities dominate and proposing a smoothing mechanism to enhance continuity and empirical performance in language and vision tasks. AI
IMPACT These studies offer theoretical insights into MoE model behavior, potentially leading to more robust and accurate AI systems in language and vision tasks.
RANK_REASON The cluster contains two academic papers published on arXiv discussing theoretical aspects of Mixture-of-Experts models.
- arXiv
- diffusion process
- Hugging Face
- Language Models
- Sparse Mixture of Experts
- Vision Models
- alphaXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- fontange
- Gotit.pub
- Influence Flower
- mixture of experts
- ScienceCast
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →