Researchers have introduced BoT-GRPO, a novel reinforcement learning algorithm designed to enhance reasoning capabilities in large language models. This method, an extension of GRPO, efficiently uses token-level rewards by aggregating them into a "bag of tokens" and calculating advantages relative to group statistics. BoT-GRPO operates without a critic, making it a direct replacement for GRPO when token-level rewards are available. Experiments show it achieves higher compile rates and faster convergence on code generation tasks compared to existing GRPO variants, and it also improves mathematical reasoning performance. AI
IMPACT This new RL algorithm could accelerate LLM reasoning and improve performance on complex tasks like code generation and mathematical problem-solving.
RANK_REASON The cluster describes a new algorithm presented in an arXiv paper. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Bag-of-Tokens Group Relative Policy Optimization
- BoT-GRPO
- GRPO
- Phi-4-mini-reasoning
- Qwen2.5-3B
- React
- SmolLM3 3B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →