Researchers have developed a new method called CHERRY for training compute-efficient language models, focusing on three key techniques. The first involves selective supervision, concentrating training on approximately 15% of output tokens that carry the most semantic meaning, which yields a 4.5x efficiency gain per supervised token. The second technique compresses a large transformer model into a smaller one by averaging layers and then restores its performance through learned recurrent unrolling, achieving a 2.5x parameter reduction. Finally, these compressed models are combined into a Mixture of Efficient Experts (MoEE) to improve performance beyond individual experts, as demonstrated on the Korean foundation model CHERRY-1.8B. AI
IMPACT Introduces novel techniques for creating more efficient language models, potentially reducing computational costs for training and inference.
RANK_REASON The cluster contains a research paper detailing novel techniques for training language models.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →