发表机构
Shanghai Jiao Tong University; Taobao & Tmall group of Alibaba(上海交通大学; 阿里巴巴淘宝天猫集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出REST框架,将RL与蒸馏联合训练,通过AMD优化,实现少步无CFG图像生成,性能优于40步RL教师,成本低、迭代少。
AI 中文摘要
高效的文本到图像生成既需要基于强化学习(RL)的奖励对齐,也需要少步蒸馏,但这些流程通常按顺序执行,会增加训练成本并存在压缩过程中奖励增益丢失的风险。我们转而采用原生RL视角:扩散RL已生成带奖励分数的有限步轨迹,其中间状态可作为蒸馏监督的天然来源,而非采样的一次性副产品。基于此见解,我们提出REST(奖励增强型带分数轨迹蒸馏),这是一种单阶段RL-蒸馏联合训练框架,可将解耦的学生模型附加到任意RL教师模型上。学生模型从教师模型不断演进的rollout轨迹中按片段学习,同时不改变教师模型的原始优化。为防止均匀模仿保留不良的低奖励行为,我们进一步引入优势调制蒸馏(AMD),将rollout优势转换为基础蒸馏损失的带符号权重。AMD强化来自偏好轨迹的监督,并适度排斥学生模型学习低奖励轨迹。所得框架轻量且即插即用,无需额外图像rollout、无需单独的蒸馏数据集、也无需对抗训练。在组合生成、视觉文本渲染和人类偏好对齐上的实验表明,REST可实现无分类器引导(CFG)的少步推理,其性能与40步RL教师模型相当或更优,且相较于纯RL的额外总训练成本低于25%。REST在DrawBench PickScore上较RTDMD提升0.82,同时仅需五分之一的训练迭代次数。
英文摘要
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide distillation supervision. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage co-training framework in which a decoupled student learns from the evolving RL teacher's trajectories without changing teacher optimization. Advantage-Modulated Distillation (AMD) transforms rollout advantages into signed weights, strengthening imitation of preferred trajectories and aligning distillation priorities with task value. The resulting framework is general and lightweight, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment demonstrate competitive few-step, CFG-free generation with RAM or DiffusionNFT teachers. With only four sampling steps, REST-RAM achieves a DrawBench PickScore of 23.97, outperforming both the 40-step RAM teacher (23.95) and RTDMD (23.71).