TDM-R1: 通过非可微奖励强化少步扩散模型
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
- Hong Kong University of Science and Technology(香港科技大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- hi-Lab, Xiaohongshu Inc(小红书实验室)
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TDM-R1通过非可微奖励强化少步扩散模型,提升文本到图像生成性能
AI中文摘要:
尽管少步生成模型在显著降低成本的情况下实现了强大的图像和视频生成能力,但针对少步模型的通用强化学习(RL)范式仍是一个未解决的问题。现有的针对少步扩散模型的RL方法强烈依赖于通过可微奖励模型进行反向传播,从而排除了大多数重要的现实世界奖励信号,例如非可微奖励,如人类的二元喜爱度、物体计数等。为了正确地将非可微奖励纳入以提高少步生成模型,我们引入了TDM-R1,这是一种基于领先少步模型轨迹分布匹配(TDM)的新强化学习范式。TDM-R1将学习过程分解为代理奖励学习和生成器学习。此外,我们开发了实际方法来获取TDM确定性生成轨迹上的每一步奖励信号,从而得到一种统一的RL后训练方法,显著提高了少步模型在通用奖励下的能力。我们进行了从文本渲染、视觉质量和偏好对齐的广泛实验。所有结果都表明,TDM-R1是少步文本到图像模型的强大强化学习范式,在域内和域外指标上都实现了最先进的强化学习性能。此外,TDM-R1也能有效扩展到最近的强Z-Image模型,始终在仅4次NFE的情况下优于其100-NFE和少步变体。项目页面:https://github.com/Luo-Yihong/TDM-R1
英文摘要:
While few-step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few-step models remain an unsolved problem. Existing RL approaches for few-step diffusion models strongly rely on back-propagating through differentiable reward models, thereby excluding the majority of important real-world reward signals, e.g., non-differentiable rewards such as humans' binary likeness, object counts, etc. To properly incorporate non-differentiable rewards to improve few-step generative models, we introduce TDM-R1, a novel reinforcement learning paradigm built upon a leading few-step model, Trajectory Distribution Matching (TDM). TDM-R1 decouples the learning process into surrogate reward learning and generator learning. Furthermore, we developed practical methods to obtain per-step reward signals along the deterministic generation trajectory of TDM, resulting in a unified RL post-training method that significantly improves few-step models' ability with generic rewards. We conduct extensive experiments ranging from text-rendering, visual quality, and preference alignment. All results demonstrate that TDM-R1 is a powerful reinforcement learning paradigm for few-step text-to-image models, achieving state-of-the-art reinforcement learning performances on both in-domain and out-of-domain metrics. Furthermore, TDM-R1 also scales effectively to the recent strong Z-Image model, consistently outperforming both its 100-NFE and few-step variants with only 4 NFEs. Project page: https://github.com/Luo-Yihong/TDM-R1