Researchers have developed ReCo, a novel reweighting method designed to improve Group Relative Policy Optimization (GRPO) in language models. GRPO, a standard reinforcement learning technique, has been observed to sometimes reduce a model's reasoning capacity by concentrating on high-probability responses. ReCo addresses this by normalizing response contributions based on their expected occurrence and replacing the token-level importance ratio with a variance-based one. This approach aims to enhance coverage of reasoning paths, particularly for larger values of k in benchmarks like Pass@k. AI
IMPACT This research could lead to language models with improved reasoning capabilities and broader coverage of potential solutions.
RANK_REASON The cluster contains a research paper detailing a new method for improving reinforcement learning in language models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DagsHub
- Group Relative Policy Optimization
- GRPO
- Hugging Face
- IArxiv
- Llama 3.1 8B-Instruct
- Qwen2.5-Math-1.5B
- Qwen2.5-Math-7B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →