发表机构
University of Sydney; Shanghai AI Laboratory; Hong Kong University of Science and Technology; Chinese University of Hong Kong(悉尼大学; 上海人工智能实验室; 香港科技大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VDOT++提出非平衡最优传输蒸馏框架,统一处理T2V、I2V和条件生成任务,通过非对称质量匹配与加权中位数聚合,实现四步生成与多步教师模型性能相当。
AI 中文摘要
视频创作涵盖文本到视频(T2V)、图像到视频(I2V)以及基于条件的生成,然而视频扩散模型仍然成本高昂,因为它们在采样过程中需要反复评估大型骨干网络。分布匹配蒸馏(DMD)降低了这一成本,但其反向Kullback-Leibler(KL)目标在学生分布与教师分布重叠有限时,可能提供不稳定或不完整的指导。VDOT通过添加最优传输蒸馏(OTD)解决了该问题,其显式耦合为基于条件的生成提供了几何方向。然而,平衡OTD在每对对应的学生-教师帧对的空间令牌之间执行全质量匹配。这一假设在T2V和I2V任务中会被削弱,因为一个条件可对应多种有效输出,且不同实现之间的空间内容无需对齐。我们提出了VDOT++,一个统一的蒸馏框架,将相同的训练方案分别应用于三个任务族的生成器。它通过非对称的非平衡公式使OTD对输出多样性具有鲁棒性,该公式允许不可靠的学生令牌携带更少质量,同时保持对教师令牌的覆盖。此外,$\nell_1$基础成本将基于均值的聚合替换为更保留模式的加权中位数,从而限制了远处传输目标的影响。这两个变化分别决定了匹配对象以及所选目标的聚合方式。我们还通过顺序反向传播结合了分布匹配与对抗性细化,并利用解耦的分数网络进行跨尺度蒸馏,其中更大的分数网络可改进紧凑生成器。在UVCBench、VBench、VBench-I2V和VACE基准上的实验表明,所得到的四步生成器在三个任务族中均与多步教师模型和强少步基线具有竞争力。
英文摘要
Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can provide unstable or incomplete guidance when the student and teacher distributions have limited overlap. VDOT addressed this issue by adding optimal transport distillation (OTD), whose explicit coupling supplies geometric directions for condition-based generation. Balanced OTD, however, performs full-mass matching between the spatial tokens of each corresponding student--teacher frame pair. This assumption weakens for T2V and I2V, where one condition admits many valid outputs and spatial content need not align across different realizations. We present VDOT++, a unified distillation framework that applies the same training recipe separately to generators for the three task families. It makes OTD robust to output diversity through an asymmetric unbalanced formulation that allows unreliable student tokens to carry less mass while maintaining coverage of the teacher tokens. An $\ell_1$ ground cost further replaces mean-based aggregation with a more mode-preserving weighted median that limits the influence of distant transport targets. The two changes respectively determine whom to match and how the selected targets should be aggregated. We additionally combine distribution matching and adversarial refinement through sequential backward passes, and exploit the decoupled score networks for cross-scale distillation, where larger score networks improve a compact generator. Experiments on UVCBench, VBench, VBench-I2V, and the VACE benchmark show that the resulting four-step generators are competitive with many-step teachers and strong few-step baselines across all three task families.