arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09123cs.CV

掩码强制:通过双噪声掩码展开改进自回归视频扩散蒸馏

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

  • HKUST(GZ)(香港科技大学(广州))
  • HKUST(香港科技大学)
  • LIGHTSPEED
  • UCSD(加利福尼亚大学圣地亚哥分校)
  • CUHK(SZ)(香港中文大学(深圳))
  • NUS(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao

AI总结:

针对自回归视频扩散蒸馏中的模式坍缩问题,提出双噪声掩码展开策略,通过随机掩码注入干净信号扰动学生自展开,提升多种蒸馏方法的视觉质量。

AI中文摘要:

自回归(AR)视频扩散模型在实时视频生成方面展现出巨大潜力。近期方法通过分布匹配蒸馏(DMD)将预训练的双向视频扩散模型蒸馏为因果自回归学生模型,但生成的视频常存在过饱和和过平滑问题,导致视觉质量和真实感有限。关键因素在于DMD中反向KL目标函数的模式寻求行为,这可能使学生分布坍缩到教师分布的少数几个模式上。为解决此问题,我们提出掩码强制(Mask Forcing),一种双噪声掩码展开策略,通过扰动自回归学生自展开过程来缓解反向KL模式寻求引起的模式坍缩。核心思想是在自回归扩散蒸馏的自展开过程中,沿空间和时间轴通过随机掩码向噪声展开输入注入更干净的信号。这种扰动鼓励学生展开探索教师分布的更多区域,使DMD能够提供超出学生已覆盖模式的学习信号。此外,更干净的令牌可作为其他更嘈杂令牌的去噪引导,改善中间展开预测并减少误差累积。大量实验表明,我们的方法在无需真实视频数据或额外后训练阶段的情况下,高效提升了多种自回归视频扩散蒸馏方法的视觉质量。

英文摘要:

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.

补充信息

↑