Hugging Face has released Olmo-core 3, an open-source framework designed to scale the training of Mixture-of-Experts (MoE) large language models to trillion-parameter sizes. This new system addresses the computational inefficiencies and high costs associated with training large MoEs by optimizing expert routing and distribution across GPUs. Benchmarks show Olmo-core 3 significantly increases training throughput, achieving up to 2.7 times faster performance compared to previous implementations, while maintaining computational efficiency even with a large number of experts. AI
IMPACT Enables more efficient training of trillion-parameter MoE models, potentially lowering costs and increasing accessibility for researchers.
RANK_REASON Open-source release of a new training infrastructure framework for large MoE models by a major AI lab (Hugging Face). [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →