DistillAlign:自回归视频蒸馏中覆盖模式与寻找模式的协调
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
- Riemann Dynamics(黎曼动力学)
- Nanyang Technological University(南洋理工大学)
- Wellington College, UK(惠灵顿公学(英国))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对自回归视频蒸馏中初始化与DMD阶段解耦、覆盖度不足等问题,提出DistillAlign联合蒸馏方法,结合模式寻找与覆盖约束,提升了视频生成的质量、覆盖度与多样性。
AI中文摘要:
现有的自回归视频蒸馏方法通常采用基于分布匹配蒸馏(Distribution Matching Distillation, DMD)的多阶段流程,但它们通常将初始化阶段与DMD阶段解耦,导致两个阶段追求不同的目标分布,且主要通过VBench等视觉分数判断中间学生模型的优劣。本文从分布视角重新审视该设计,鉴于分布匹配损失的模式寻找特性,良好的初始化应匹配目标DMD教师的模式覆盖度,而非仅追求高质量。为分析这一点,我们引入了一种分布评估协议,在共享潜在空间中测量学生与教师分布的精度和覆盖度,该协议能揭示视觉分数隐藏的差异:部分初始化达到高精度但低覆盖度,导致优化效果不佳,而模式覆盖型初始化能保留更广泛的支撑集。此外,即使目标分布已对齐,DMD的反向KL目标在训练后期仍会驱动学生向教师的高概率区域靠近,从而降低覆盖度和多样性。为解决这一问题,我们提出了联合蒸馏方法,将DMD的模式寻找目标与基于一致性蒸馏的模式覆盖约束相结合。实验表明,我们的方法提升了生成质量、覆盖度和多样性;值得注意的是,即使使用Wan-1.3B DMD教师,它仍优于使用Wan-14B优化的基线,凸显了分布对齐在自回归视频蒸馏中的重要性。
英文摘要:
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.