发表机构
Peking University; Nanjing University; Stanford University(北京大学; 南京大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型RL偏好对齐的奖励不匹配问题,提出SGPO方法,通过分阶段分配目标,使生成质量提升26.7%、收敛速度提高36.7%。
AI 中文摘要
扩散模型具有强大的生成能力,但其最大似然训练目标仅聚焦于重构数据分布,难以与特定偏好对齐。用于扩散模型偏好对齐的强化学习(RL)很有前景,但受限于奖励稀疏性:单个奖励无法支撑优化,现有RL方法通常将最终奖励反向传播至所有前序步骤。然而,去噪是分阶段进行的,具有不同的语义和可控性,在所有步骤中重复最终奖励会产生时间目标不匹配,催生奖励捷径,进而导致奖励黑客行为;同时,由于奖励回填,每个时间步收到相同奖励,无法区分动作,削弱了优化过程。为解决该问题,我们提出面向扩散模型的阶段引导式逐步优化(SGPO)方法,该方法联合利用信噪比和语义变化识别生成阶段,并自适应分配阶段特定目标:早期去噪处于混沌状态,远离最终奖励,奖励-行为相关性弱,此阶段应优先退出混沌状态;中期潜在变量过渡到稳定结构,最终奖励与生成行为更匹配,因此该阶段在优化最终奖励的同时探索多样性,避免过早收敛到单一模式;后期潜在变量核心结构基本固定,偏好优化主要放大局部细节,存在过拟合风险,因此优先稳定收敛以避免质量下降。16项对比实验结果验证了SGPO的有效性:我们的方法在生成质量上实现了26.7%的平均提升,收敛速度提高了36.7%。
英文摘要
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences. Reinforcement learning (RL) for preference alignment in diffusion models is promising but limited by reward sparsity. Since a single reward cannot support optimization, existing RL methods usually backpropagate the final reward to all previous steps. However, denoising is stage-wise, with distinct semantics and controllability. Repeating the final reward across all steps creates a temporal objective mismatch, encouraging reward shortcuts that lead to reward hacking. At the same time, due to reward backfilling, each time step receives the same reward, making it impossible to distinguish between actions, thereby weakening the optimization process. To resolve this issue, we propose Stage-Guided Per-Step Optimization (SGPO) for diffusion models, which jointly leverages signal-to-noise ratio and semantic changes to identify generation stages and adaptively assign stage-specific objectives. Early denoising is chaotic and far from the final reward, resulting in weak reward-behavior correlation. This stage should prioritize exiting the chaotic state. In the mid stage, the latent transitions to a stable structure, where the final reward better corresponds to generative behavior. Therefore, this stage optimizes the final reward while exploring diversity to avoid early convergence to a single mode. In the late stage, the latent's core structure is largely fixed, and preference optimization mainly amplifies local details, risking overfitting. Therefore, stable convergence is preferred to avoid quality degradation. Results from 16 comparative experiments validate SGPO. Our method achieves 26.7% average gains in generative quality and 36.7% higher convergence speed.