Researchers have developed RAPTOR, a novel framework for differentially private training of Mixture-of-Experts (MoE) models. Existing methods treat these sparse models as dense blocks, leading to issues like gradient suppression and diluted updates. RAPTOR addresses these by alternating shared and expert optimization, employing expert-specific clipping and noise, and a privacy-free rule for selecting layers to protect from routing entropy. Experiments on models like Switch Transformer and OLMoE demonstrate consistent performance gains over standard DP baselines, particularly at tighter privacy budgets. AI
IMPACT Enhances privacy guarantees for large, sparse AI models, potentially enabling wider adoption in sensitive applications.
RANK_REASON The cluster describes a new research paper introducing a novel training framework for AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- DeepSeek-VL2-Tiny
- Glue
- Hugging Face
- mixture of experts
- Olmoe
- RAPTOR
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →