Researchers have developed ScalePRM, a novel method for training process reward models (PRMs) that bypasses the need for expensive step-level correctness labels or ground-truth answers. This approach scales verification compute by generating and aggregating multiple independent verifications of each reasoning step to create synthetic labels. ScalePRM achieved a 67.5 F1 score on the ProcessBench benchmark, outperforming traditional ground-truth-based training and even GPT-4o as a critic. When used as a reward signal for RL training with Qwen2.5-Math-7B, it improved average accuracy across six mathematical reasoning benchmarks to 47.4%, surpassing ground-truth-based RLVR. AI
IMPACT This method could significantly reduce the cost and complexity of training AI models that require step-by-step reasoning, potentially accelerating progress in areas like mathematical problem-solving.
RANK_REASON The cluster describes a new research paper detailing a novel method for training AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →