Researchers have developed a new actor-critic algorithm called AC2, designed to improve the efficiency of training large language models (LLMs) with reinforcement learning. Unlike standard methods that require rolling out trajectories to their terminal reward, AC2 assigns credit to "action chunks" and uses a learned critic to score these segments. This approach allows the policy to update without observing a full terminal reward, significantly reducing computational costs. The method was tested on Qwen3-4B and demonstrated superior performance on the IMO-ProofBench benchmark compared to GRPO, achieving a higher validation score with substantially fewer decoding FLOPs and training steps. AI
IMPACT This new algorithm could significantly reduce the computational cost and time required for training large language models, potentially accelerating research and development in the field.
RANK_REASON Academic paper detailing a new algorithm for LLM training. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →