Researchers have developed a novel method for training language models using reinforcement learning with verifiable rewards (RLVR) while adhering to prompt-level differential privacy. This approach ensures that the released model weights are differentially private with respect to individual training problems. The method aggregates gradients, clips contributions, adds Gaussian noise, and composes privacy loss, providing the first known differential privacy guarantee for RLVR training. Experiments with Qwen2.5-1.5B-Instruct demonstrated that the reward signal significantly improves accuracy on mathematical tasks like MATH and GSM8K, even under privacy constraints, outperforming supervised fine-tuning recipes and retaining most of the gains from non-private methods. AI
IMPACT This research could enable the development of more private AI models, particularly for sensitive data applications, without significant performance degradation.
RANK_REASON The cluster contains an academic paper detailing a new method for training language models with differential privacy guarantees. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →