Researchers have introduced Mixture of Reused Experts (MoRE), a novel neural network architecture that combines the parameter efficiency of Recurrent Transformers with the capacity of Mixture-of-Experts (MoE) models. MoRE addresses the high memory footprint of traditional MoEs by sharing expert pools across adjacent layers, allowing for greater routing diversity without increasing parameters. Experiments demonstrate that MoRE achieves superior performance in perplexity and downstream tasks compared to existing weight-sharing architectures and standard MoEs, across various model scales. AI
IMPACT This architecture could lead to more parameter-efficient large language models, reducing memory requirements for training and inference.
RANK_REASON The cluster contains a research paper detailing a new model architecture. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →