发表机构
Westlake University; Zhejiang University; Kuaishou Technology(西湖大学; 浙江大学; 快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlowCPO提出统一散度框架,通过离线前向KL目标同时利用偏好与非偏好样本,无需在线采样,在域内任务上优于FlowDPO等基线。
AI 中文摘要
流模型和扩散模型的偏好对齐现在涵盖了在线强化学习和离线偏好优化,但这些方法之间的关系仍不清楚。特别是,现有的前向过程对齐方法需要从当前模型获取新样本,而基于固定偏好对的离线方法主要依赖于仅正向微调或DPO风格似然比替代目标。我们通过一个基于散度的框架来组织这些方法,并引入FlowCPO,一种离线前向KL目标,它同时使用偏好和非偏好样本,无需在线采样。对于线性插值,我们在明确的规则性条件下证明前向KL目标受对比流匹配损失约束,从而在固定数据上产生一个可处理的替代目标。我们进一步证明该损失是非负的,而简化FlowDPO的有符号回归损失可能无下界。在域内设置中,FlowCPO在GenEval和OCR平均得分上优于评估的基线,在CFG 3.0下分别达到0.84和0.87,而FlowDPO为0.81和0.74。在域外设置中,结果好坏参半,FlowCPO在GenEval上取得最佳结果,但在多个指标上的奖励得分低于RFT。
英文摘要
Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on fixed preference pairs rely primarily on positive-only fine-tuning or DPO-style likelihood-ratio surrogates. We organize these approaches through a divergence-based framework and introduce FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts. For linear interpolation, we show under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data. We further show that this loss is nonnegative, whereas the signed regression loss of simplified FlowDPO can be unbounded below. In the in-domain setting, FlowCPO achieves higher mean GenEval and OCR scores than the evaluated baselines, reaching 0.84 and 0.87 versus 0.81 and 0.74 for FlowDPO at CFG 3.0. In the out-of-domain setting, the results are mixed, with the best GenEval result but lower reward scores than RFT on several metrics.