PulseAugur
实时 15:12:05
English(EN) Normalized Rewards for Preference Optimization

新的正则化技术提高了大型语言模型的对齐度和基准性能

研究人员为直接对齐算法(DAA),如DPO,开发了一种新的正则化技术,以减轻大型语言模型(LLMs)的过度优化问题。该方法旨在维持首选响应的似然性,从而改善生成质量与通用基准能力之间的权衡。该正则化技术应用于基于参考和无参考的方法,在AlpacaEval2等基准测试和通用性能指标上均有所提升,特别是对于Llama 3.1 8B-Instruct等模型。 AI

影响 这项研究通过改进对齐技术,有望带来更稳定、更强大的大型语言模型,从而提升在各种基准测试上的性能。

排序理由 该集群包含两篇学术论文,讨论了在机器学习环境中优化奖励函数的方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的正则化技术提高了大型语言模型的对齐度和基准性能

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf ·

    偏好优化中的归一化奖励

    arXiv:2607.16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood…

  2. arXiv cs.LG TIER_1 English(EN) · Francesco Bacchiocchi, Matteo Castiglioni, Alberto Marchesi, Nicola Gatti ·

    为分段线性奖励最小化遗憾:合同、拍卖及其他

    arXiv:2503.01701v2 Announce Type: replace-cross Abstract: Most microeconomic models of interest involve optimizing a piecewise linear function. These include contract design in hidden-action principal-agent problems, selling an item in posted-price auctions, and bidding in first-…