PulseAugur
实时 05:21:37
English(EN) Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

新的ThermoDPO方法解决了生成模型对齐中的流形漂移问题

研究人员发现“流形漂移”是将其偏好优化方法扩展到生成模型的连续时间动力学中的一个关键问题。当流匹配中的奖励驱动更新将终端样本移出预训练数据流形时,就会发生这种现象。为了解决这个问题,提出了一种名为ThermoDPO的新方法,该方法使用温度控制将成对偏好优化锚定在首选样本上。加权变体ThermoDPO-weighted进一步提高了性能,在基准测试中达到了0.899的StrictScore,并在SD3.5-M模型上的OCR和其他指标上取得了显著的提升。 AI

影响 引入了一种新颖的技术来提高生成模型的对齐和稳定性,有望带来更可靠、更准确的输出。

排序理由 详细介绍生成模型对齐新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的ThermoDPO方法解决了生成模型对齐中的流形漂移问题

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yansen Han, Shengyi Liao, Yuanxing Zhang, Pengfei Wan, Tao Lin ·

    流偏好优化中的流形漂移:奖励破解的根本原因

    arXiv:2608.20011v1 Announce Type: new Abstract: Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inheren…