为扩散模型设计强化学习:统一的路径空间视角
Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
- State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
- School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文从路径空间视角统一了扩散模型RL算法的原理,推导得到降方差值梯度形式,提出多样本KDE估计器和尺度受限权重族,在SD3.5-M等模型上验证了方法有效性并优于基线。
AI中文摘要:
强化学习(RL)后训练为让扩散模型对齐人类偏好和特定任务奖励提供了直接途径,但当前适用于扩散模型的RL算法仍零散分散:反向轨迹方法依赖离散化似然比,而前向匹配方法则基于带奖励标签的回退样本加噪版本进行训练。本文表明这些看似不同的损失源于单一的路径空间原理。从正则化扩散-RL目标出发,我们利用采样SDE之间的重要性采样得到轨迹空间上的显式策略梯度估计器,该估计器包含Flow-GRPO型更新所隐含的随机伊藤积分;我们推导了等效的降方差值梯度形式,该形式可恢复AWM和DiffusionNFT的前向匹配结构。这明确了这些方法族之间的经验差距源于降方差效应,而非RL原理的差异。该推导产生了由值梯度估计、权重函数和采样选择组织的统一设计空间。在该空间内,我们提出了一种多样本KDE值梯度估计器,该估计器复用回退组,同时结合尺度受限的权重族,在保留现有稳定方案的同时排除奇异方案。在SD3.5-M和Qwen-Image模型上进行的实验验证了降方差解释,并表明所得方案优于现有扩散-RL基线。
英文摘要:
Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic Itô integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe improves over prior diffusion-RL baselines.