arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过速度匹配扩展扩散模型的强化学习

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

Jaemoo Choi, Wei Guo, Yuchen Zhu, Arash Vahdat, Molei Tao, Julius Berner, Yongxin Chen

arXiv 2608.23664首次发表:更新:

发表机构

Georgia Institute of Technology; NVIDIA(佐治亚理工学院; 英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于奖励的速度匹配(RVM)方法,用于扩散模型的奖励微调,其训练成本显著降低,性能优于或相当基于轨迹的策略梯度方法,还通过动态跟踪奖励改善了视频生成的运动效果。

AI 中文摘要

奖励微调正成为使扩散模型适应人类偏好和特定任务目标的重要工具,但现有方法大多继承了大型语言模型的策略梯度机制。与自回归模型不同,扩散模型无法为生成样本提供易处理的似然度。因此,当前方法要么从随机去噪转移构造轨迹似然度,要么用证据下界近似端点似然度,这会引入额外的计算和算法复杂度。我们证明,这种基于似然的机制对于有效的扩散模型奖励微调并非必需。我们提出了基于奖励的速度匹配(RVM),这是一种简单的无轨迹更新,直接作用于速度场。RVM强化与高奖励生成相关的方向,抑制与低奖励相关的方向,并包含一个可选的锚定项以控制与参考速度的漂移。值得注意的是,它提供了一个通用框架,可将近期的微调方法(包括RAM和DiffusionNFT)作为特例恢复。在各种大规模扩散模型奖励微调任务中,RVM在训练成本显著降低的情况下,性能与基于轨迹的策略梯度方法相当或更优。我们进一步发现,一旦速度更新被简化,特定的损失变体的重要性就低于奖励和锚定设计。对于视频生成,标准偏好奖励可能倾向于视觉干净但几乎静态的输出;引入新的动态跟踪奖励可显著改善运动,同时提升整体VBench性能。这些结果表明,扩散模型的可扩展奖励微调更适合用原生速度表示来实现,而非基于似然的策略优化。

英文摘要

Reward-based fine-tuning of diffusion models has largely inherited likelihood-based policy optimization developed for autoregressive large language models (LLMs). Diffusion models, however, are natively trained through velocity regression and do not directly provide the likelihood of a generated sample. Existing methods address this using transition likelihoods along stochastic denoising trajectories or evidence lower bounds (ELBOs) to approximate generated-sample likelihoods. These approximations arise from applying likelihood-based updates to models whose native training and generation operate through velocity fields. We instead take velocity matching as the starting point for reward fine-tuning. We propose \textbf{reward-based velocity matching (RVM)}, which weights the velocity-matching loss by reward with an additional anchor regression term, requiring neither likelihood estimation nor likelihood ratios. RVM recovers Reinforce Adjoint Matching at the update level and contains DiffusionNFT as a special case, while ELBO-based methods are closely related through the same velocity-regression structure. This unified view isolates reward design and anchor velocity as principal design choices in velocity-based fine-tuning. Across text-to-image, text-to-video, and image-to-video generation, RVM matches or outperforms the evaluated trajectory-based methods. On Wan2.1-T2V-1.3B, it achieves the highest VBench Overall among the evaluated methods, with substantially lower estimated training costs than the trajectory-based approaches. For video fine-tuning, we introduce a dynamic-tracking reward that provides explicit motion feedback and improves both Dynamic Degree and overall VBench performance.

Comments32 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑