Researchers have introduced GeoRA, a novel low-rank adaptation method specifically designed for Reinforcement Learning with Verifiable Rewards (RLVR). Unlike existing methods that focus on supervised fine-tuning, GeoRA accounts for the unique optimization dynamics and geometric structures inherent in RLVR. By leveraging singular value decomposition to extract principal directions of the RL update subspace, GeoRA effectively preserves pre-trained model structures while enabling efficient computation. Experiments on Qwen and Llama models demonstrate that GeoRA outperforms other low-rank adaptation techniques in various RLVR tasks, including mathematics, medicine, and coding, while also showing improved generalization and reduced forgetting on out-of-domain datasets. AI
IMPACT GeoRA's tailored approach to RLVR could improve the reasoning capabilities and efficiency of large language models across various domains.
RANK_REASON The cluster contains a research paper detailing a new method for adapting large language models. [lever_c_demoted from research: ic=1 ai=1.0]
- Jiaying Zhang
- Llama
- Qwen
- Reinforcement Learning with Verifiable Rewards
- singular value decomposition
- supervised fine-tuning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →