Researchers have identified a misalignment between standard cross-entropy (CE) training and the objective of producing correct outputs in verifiable domains like mathematical reasoning and code generation. This issue arises because CE training can inadvertently assign higher probabilities to incorrect outputs, even when it accurately imitates expert demonstrations. To address this, a new method called entropy-regularized cross-entropy (ER-CE) is proposed, which uses token-level Shannon entropy as a proxy to control the policy's support and prevent mass from spreading to unsupported outputs. Experiments on mathematical reasoning and code-generation benchmarks show that ER-CE consistently improves verifier accuracy compared to standard CE. AI
IMPACT This research offers a practical method to improve the accuracy of AI models in tasks requiring verifiable outputs, potentially leading to more reliable AI systems in fields like coding and mathematics.
RANK_REASON The cluster contains an academic paper detailing a new method for training AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →