发表机构
Sungkyunkwan University; Korea Institute for Advanced Study(成均馆大学; 韩国高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单步生成模型奖励引导微调未被充分探索的问题,本文从最优传输视角研究Wasserstein梯度流,提出无需奖励梯度的新型奖励引导微调方法,在多数据集多奖励任务上实现了更优奖励对齐效果。
AI 中文摘要
为降低生成模型的时间复杂度,单步生成模型近期通过从噪声到数据的单次前向传播直接映射的方式出现,但单步生成模型的奖励引导微调方法仍未被充分探索。为解决该问题,我们从最优传输视角研究单步生成器,探讨Wasserstein梯度流(WGF)以建模概率空间中平滑且可控的分布演化,进而提出一种通过WGF实现的单步生成模型的新型奖励引导微调方法。我们推导了一种无需奖励梯度的实用训练方法,可同时处理不可微和可微奖励;此外,该方法能提供平滑稳定的奖励引导分布更新,同时缓解奖励黑客行为和模式崩溃。在2D合成数据、CIFAR-10、256×256 ImageNet上,针对JPEG(非)压缩性、类别概率、黑白、CLIP对齐等多样奖励的实验表明,与基线方法相比,我们的方法实现了更优的奖励对齐效果。
英文摘要
To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a single forward pass. However, the reward-guided fine-tuning method of one-step generative models remains largely unexplored. To address this, we consider one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF) for modeling smooth and controlled distributional evolution in probability space. We then propose a novel reward-guided fine-tuning of a one-step generative model via WGF. We derive a practical training method that requires no reward gradients, thereby handling both non-differentiable and differentiable rewards. Moreover, our method provides smooth and stable reward-guided distributional updates while mitigating reward hacking and mode collapse. Experiments on 2D synthetic data, CIFAR-10, and ImageNet 256$\times$256 with diverse rewards, including JPEG (in)compressibility, class probability, Black-and-White and CLIP alignment, show that our method achieves better reward alignment compared to baselines.
Comments14 pages, 9 figures