arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20011cs.AIcs.CV

流偏好优化中的流形漂移:奖励黑客的根本原因

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

  • Zhejiang University(浙江大学)
  • Kuaishou Technology(快手科技)
  • Westlake University(西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

Yansen Han, Shengyi Liao, Yuanxing Zhang, Pengfei Wan, Tao Lin

AI总结:

本文针对流偏好优化中流形漂移导致奖励黑客的问题,提出ThermoDPO及其加权变体,在玩具基准和SD3.5-M数据集上均取得优于FlowDPO等方法的性能。

AI中文摘要:

偏好优化是生成式模型的标准对齐方法,但将其扩展到连续时间动力学仍非易事。在流匹配(flow matching)中,奖励驱动的更新会修改传输轨迹,且不存在对预训练数据流形的固有约束,可能会将终端样本移出预训练支撑集。我们将这种失效模式形式化为流形漂移(manifold drift)。理论上,我们证明最优流匹配可恢复终端数据分布,而偏好更新只要其诱导的终端位移具有非零法向分量,就会使预训练流形偏离。作为解决方案,我们提出ThermoDPO,这是一种温度控制的目标函数,可将成对偏好优化锚定在偏好样本上。在不同温度区间,该目标函数连接拒绝采样微调(rejection sampling fine-tuning)和FlowDPO,并控制基于逐点重建的流形距离替代物。为应对低温下信号减弱的问题,我们进一步提出加权变体ThermoDPO-weighted。在主要玩具基准上,ThermoDPO-weighted的StrictScore达到0.899,而FlowDPO为0.629,FlowDPO+RFT为0.857;在CFG=4.5的SD3.5-M上,它将OCR提升47.5%,四个指标的平均值提升16.0%。

英文摘要:

Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretrained support. We formalize this failure mode as manifold drift. Theoretically, we show that optimal flow matching recovers the terminal data distribution, whereas a preference update leaves the pretrained manifold whenever its induced terminal displacement has a nonzero normal component. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO and controls a pointwise reconstruction-based surrogate for manifold distance. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted. On the main toy benchmark, ThermoDPO-weighted attains a StrictScore of 0.899, compared with 0.629 for FlowDPO and 0.857 for FlowDPO+RFT. On SD3.5-M at CFG = 4.5, it improves OCR by 47.5% and the average of four metrics by 16.0%.

↑