English(EN)BiasReducer: Adaptive Bias Mitigation for Reward Models
新研究解决AI中奖励模型的不稳定性和效率问题
作者PulseAugur 编辑部·[8 个来源]·
研究人员正在开发新方法来改进用于大型语言模型和扩散模型中的奖励模型的训练和鲁棒性。一种名为“密度感知奖励聚合”(DARA)的方法,通过根据奖励的活动密度对其进行加权来解决多奖励强化学习中不均衡的学习进展问题,从而在工具调用和数学推理等任务上实现更快的收敛。另一个研究重点是缓解奖励模型中的偏好不稳定性,例如使用稀疏自编码器(SAEs)识别和抑制导致矛盾判断的脆弱特征。针对视频理解,已创建了一个新的基准(VURB)和数据集(VUP-35K)来促进更鲁棒的奖励模型的开发,而CoRe和扩散奖励模型(DRM)等技术分别旨在提高生成质量和处理多模态奖励结构。
AI
arXiv:2610.00574v1 Announce Type: new Abstract: Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout gro…
arXiv cs.LG
TIER_1English(EN)·Shunchang Liu, Xin Chen, Belen Martin Urcelay, Francesco Croce·
arXiv:2605.16339v2 Announce Type: replace Abstract: Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contradictory preference assignments in response to s…
arXiv:2605.07872v2 Announce Type: replace-cross Abstract: Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality pref…
arXiv cs.AI
TIER_1English(EN)·Zhaolong Su, Yujin Han, Feng Wang, Jameson Dong, Hins Hu, Difan Zou·
arXiv:2609.36245v1 Announce Type: new Abstract: Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward …
arXiv:2609.33803v2 Announce Type: replace-cross Abstract: Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human …
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit …
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing m…
<!-- SC_OFF --><div class="md"><p>I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run.</p> <p>I find it rather funny that most of the tasks …