arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEP-OPD:重新思考扩散模型在线策略蒸馏中的输出目标与内部动态

STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models

Qingyan Wei, Guangzhao Li, Xiaobing Tu, Yinggui Wang, Xiantao Zhang, Jinkui Ren, Xiaohong Liu, Linfeng Zhang

arXiv 2608.04887首次发表:更新:

AI 中文总结

该研究针对扩散模型在线策略蒸馏提出STEP-OPD框架,通过扩展学习目标、对齐表示变化,提升了生成模型性能,在多项任务中优于标准OPD方法及对应单任务教师模型。

AI 中文摘要

在线策略蒸馏(OPD)已成为将多个任务专用图像生成模型整合为单个学生模型的有效方法。然而,现有OPD方法主要优化学生模型以匹配教师模型的输出速度,使教师模型成为优化目标的上限。仅输出级别的监督会导致学生模型的分块表示演化约束不足,削弱了跨层逐步发展的能力迁移。我们提出STEP-OPD,一种用于图像生成的在线策略蒸馏框架,它将学生模型的学习目标扩展到教师模型之外,并对其内部表示演化引入显式约束。我们不将教师模型视为最终目标,而是将每个任务专用教师模型与共享基础模型之间的速度差作为进一步学习的方向,并将该差值的缩放版本添加到教师速度中。此外,我们使学生模型与教师模型之间的表示变化方向和幅度对齐,让学生模型学习表示如何跨网络块逐步变换。在组合对齐、文本渲染和人类偏好方面的实验表明,我们的方法始终优于标准OPD方法。特别是,它将DiffusionOPD的GenEval分数从0.927提高到0.961,同时还提高了OCR和所有基于偏好的指标。最终得到的统一学生模型在所有三个能力组中均超过了相应的单任务教师模型,表明输出外推实现了超越教师模型的学习,而表示变化对齐为学生模型的内部变换提供了补充指导。

英文摘要

On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD methods optimize the student mainly to match the teacher's output velocity, making the teacher the upper limit of the optimization objective. While output-level supervision alone leaves the student's blockwise representation evolution underconstrained, which weakens the transfer of capabilities that must be progressively developed across layers. We propose STEP-OPD, an on-policy distillation framework for image generation that extends the student's learning target beyond the teacher and introduces explicit constraints on its internal representation evolution. Instead of treating the teacher as the final target, we use the velocity difference between each task-specific teacher and the shared base model as a direction for further learning and add a scaled version of this difference to the teacher velocity. In addition, we align the direction and magnitude of representation changes between the student and teacher, enabling the student to learn how representations are progressively transformed across network blocks. Experiments on compositional alignment, text rendering, and human preference show that our method consistently improves Standard OPD methods. In particular, it increases the GenEval score of DiffusionOPD from 0.927 to 0.961, while also improving OCR and all preference-based metrics. The resulting unified student surpasses the corresponding single-task teachers across all three capability groups, showing that output extrapolation enables beyond-teacher learning. And representation change alignment provides complementary guidance for the student's internal transformations.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑