arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MEND:通过近端速度匹配实现流模型的强化学习

MEND: RL For Flow Models via Proximal Velocity Matching

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik

arXiv 2610.05954首次发表:更新:

发表机构

The University of Texas at Austin; Google; University of Colorado Boulder(德克萨斯大学奥斯汀分校; 谷歌; 科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MEND,一种基于近端速度匹配的流模型强化学习方法,通过奖励上限和位移代价控制更新,在少量更新下超越现有方法,适用于任意可微奖励的流模型。

AI 中文摘要

流模型的奖励后训练要么在KL惩罚或冻结参考模型下对模型自身样本重新加权,通常进行数千次更新,要么反向传播奖励并移动每个样本,而不检查该移动是否值得其大小。我们提出MEND,一种基于近端速度匹配的强化学习方法。MEND在每个提示组内设置奖励上限,因此已经得分良好的样本不会受到移动。在上限之下,它沿着奖励梯度提出移动,并且仅当有上限的奖励增益超过二次位移代价时才接受该移动。然后模型回归到由此产生的速度目标,无需KL项、冻结参考模型或优势权重。在100次更新中,MEND在六个评估器中的五个上,在与基础模型图像相同距离下,优于Flow-GRPO(约4000次更新)。在同等预算协议下,它在四种训练奖励的每次评估更新中都超过了ReFL和DiffusionNFT,分别达到PickScore 24.03对23.92和23.43。一次300次更新的三奖励运行也在其训练的所有三个奖励上超过了五奖励的DiffusionNFT模型。MEND具有通用性且易于采用:它适用于任何具有可微奖励的流骨干网络。

英文摘要

Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑