Researchers have introduced Criterion-Distilled Policy Optimization (CriPO), a novel method to enhance rubric-based Reinforcement Learning (RL) for Large Language Models (LLMs). CriPO addresses two key limitations: Unexplored Criteria (UC) where no rollout satisfies a given criterion, and Suppressed Criteria (SC) where learning signals are lost during optimization. By employing on-policy self-distillation, CriPO constructs a self-teacher to inject missing behaviors for UC and uses a counterfactual self-teacher to preserve useful patterns in negative-advantage rollouts for SC. Experiments on medicine and science benchmarks show CriPO outperforms existing rubric-based RL methods, achieving better performance with approximately half the optimization steps. AI
IMPACT This research could lead to more efficient and effective training of LLMs for complex, open-ended tasks by improving exploration and learning signal utilization.
RANK_REASON The item is an academic paper detailing a new method for enhancing RL in LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- Criterion-Distilled Policy Optimization
- Hugging Face
- KL loss
- Rubric-based RL
- Suppressed Criteria
- Unexplored Criteria
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →