Researchers have introduced Consensus Group Relative Policy Optimization (C-GRPO), a novel method for text generation that aims to reduce computational costs associated with sample-and-rerank decoding. Unlike previous approaches that often require gold references or explicit preference data, C-GRPO distills Minimum Bayes Risk (MBR) decoding into training using only a utility function and policy samples. The proposed objective function is shown to align with MBR decoding's expected-utility objective, offering a convergence guarantee. Experiments on machine translation and text summarization tasks indicate that C-GRPO achieves performance comparable to MBR decoding while outperforming other reference-free methods. AI
IMPACT This new method could lead to more efficient text generation models by reducing computational overhead during inference.
RANK_REASON The cluster contains a research paper detailing a new method for text generation. [lever_c_demoted from research: ic=1 ai=1.0]
- C-GRPO
- Consensus Group Relative Policy Optimization
- Minimum Bayes-risk automatic speech recognition
- WMT 2024
- XSum dataset
- Yuki Ichihara
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →