发表机构
ByteDance Seed; UC Berkeley(字节跳动Seed团队; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对一步生成模型微调困难的问题,提出奖励加权传输蒸馏(RWTD),通过混合倾斜当前与参考分布、特征空间最优传输和不动点回归,平衡奖励适应与知识保留,显著提升GenEval得分并保持组合能力。
AI 中文摘要
一步生成器能够通过单次网络评估实现高质量的视觉生成,但其训练后的微调(post-training)较为困难:一般的隐式生成器既不提供可处理的似然,也不提供去噪轨迹,且许多奖励函数是不可微的。我们提出了奖励加权传输蒸馏(Reward-Weighted Transport Distillation, RWTD),这是一种仅需生成样本和标量奖励评估的微调方法。RWTD并非仅对齐传统的奖励倾斜参考分布,而是构建一个自适应目标,该目标混合了分别倾斜的当前分布和参考分布。当前分量纳入了训练过程中发现的改进,而参考分量则将目标锚定在预训练生成器上。RWTD通过特征空间最优传输和不动点回归来实现这一目标。理论分析表明,RWTD的不动点分布在参考分布的离策略奖励倾斜与当前模型的在策略倾斜之间进行插值,为平衡奖励适应与保留先验知识提供了一种有原则的方法。实验上,RWTD将一步SANA Sprint 1.6B骨干网络的GenEval得分从0.73显著提升至0.80,同时,独立的偏好对齐实验展示了强大的跨奖励泛化能力,实现了均衡的提升并保持了组合能力。
英文摘要
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution, RWTD constructs an adaptive target that mixes separately tilted current and reference distributions. The current component incorporates improvements discovered during training, while the reference component anchors the target to the pretrained generator. RWTD realizes this target through feature-space optimal transport and fixed-point regression. Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge. Empirically, RWTD substantially improves the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80, while separate preference alignment experiments demonstrate strong cross-reward generalization that yields balanced improvements and preservation of compositional capabilities.