Researchers have developed UniMoMo, a framework designed to accelerate large recommendation models that use sparse mixture-of-experts (MoE) layers. This post-training compression method converts a full MoE checkpoint into a smaller, standard MoE with a reduced expert budget without requiring an additional online compression module. UniMoMo groups experts based on functional similarity, using an unlabeled calibration set to measure their response to recommendation states, and incorporates a layer-adaptive protection mechanism to safeguard high-traffic experts. Experiments on Amazon Beauty, KuaiRec, and TenRec datasets demonstrated that checkpoints with four experts achieved high NDCG@10 ratios (99.92%-102.30%) and significant speedups (1.28x-1.63x), with even more aggressive two-expert configurations showing comparable or better performance. AI
IMPACT This research could lead to more efficient deployment of large recommendation models, reducing computational costs and improving inference speed.
RANK_REASON The cluster contains a research paper detailing a new framework for model acceleration. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →