扩散模型中的在线策略自蒸馏
On-Policy Self-Distillation in Diffusion Models
AI总结:
本文提出DiffusionOPSD在线策略自蒸馏框架,将图像级奖励指导转为中间监督,在SD 3.5-M等模型上获最佳保留得分,训练GPU小时显著减少,为高效可诊断的扩散模型对齐提供新路径。
AI中文摘要:
强化学习可使扩散模型与人类偏好及特定任务目标对齐,但终点奖励未明确中间去噪预测应如何变化。我们提出DiffusionOPSD,一种在线策略自蒸馏框架,将图像级奖励指导转换为采样查询处干净输出预测的显式目标。在每次外层迭代中,冻结的行为策略生成轨迹并提供查询状态与锚点;奖励梯度在每个锚点周围构建有界正、负目标;可训练策略通过有限拟合将这些目标作为分离的监督进行拟合,之后指数移动平均更新刷新行为策略。该设置让我们可分别测量目标构建与有限实现。受控同查询实验显示,更大的目标构建增益未必在单次拟合更新后转化为更大的实现增益。在SD 3.5-M和经分步蒸馏的Z-Image-Turbo上,我们的方法在两个主干网络、10个评估器的20个奖励匹配设置中,19个取得最佳保留得分;其性能优于最强竞争方法达44.0%,且在SD 3.5-M上相较DiffusionNFT减少40%训练GPU小时,在Z-Image-Turbo上减少63%训练GPU小时。这些结果表明,在线策略自蒸馏是一种高效且可分析的扩散模型后训练方法,它将图像级奖励指导转换为显式且持续更新的中间监督,为实现更高效、可诊断的对齐开辟了路径。
英文摘要:
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.