arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定性DMD:恢复少步自回归视频蒸馏中的多样性

Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation

Zixuan Duan, Xunzhi Xiang, Yabo Chen, Xin Zhang, Changhan Liu, Haibin Huang, Chi Zhang, Qi Fan, Xuelong Li

arXiv 2609.11265首次发表:更新:

发表机构

Nanjing University; Institute of Artificial Intelligence, China Telecom (TeleAI); Fudan University(南京大学; 中国电信人工智能研究院(TeleAI); 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对少步自回归视频蒸馏中的多样性崩溃问题,提出不确定性DMD框架,通过时间步扰动和随机缓存写入恢复随机性,无需架构改动,提升多样性且保持视觉质量。

AI 中文摘要

少步蒸馏提高了自回归(AR)视频生成的效率,但常常导致多样性崩溃:在相同提示下,不同的噪声样本倾向于生成高度相似的视频,且运动动态减弱。我们分析了分布匹配蒸馏(DMD)蒸馏的AR视频生成器中的这种退化,并发现,在自回归设置中,它表现为一种结构化的不确定性崩溃:DMD的模态寻求偏差将不同的噪声样本映射到几乎相同的第一块,而确定性的AR缓存随后将这种崩溃状态传播到所有后续块,将滚动根部的局部随机性损失转变为时间变化的全局抑制。基于这一分析,我们提出了不确定性DMD,一个简单的不确定性注入框架,在AR生成的两个关键阶段恢复随机性:对第一块的时间步扰动以增加第一块的多样性,以及对后续块的随机缓存写入机制以在自回归条件中保留不确定性。该方法不需要架构更改,仅引入轻量级扰动操作。相同的扰动机制在训练和推理期间均使用。实验表明,不确定性DMD持续改善多样性和运动动态,同时保持可比的每样本视觉质量。

英文摘要

Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to produce highly similar videos with weakened motion dynamics. We analyze this degradation in Distribution Matching Distillation (DMD)-distilled AR video generators and find that, in the autoregressive setting, it takes the form of a structured uncertainty collapse: the mode-seeking bias of DMD maps different noise samples to nearly identical first chunks, and the deterministic AR cache then propagates this collapsed state to all subsequent chunks, turning a local loss of stochasticity at the rollout root into a global suppression of temporal variation. Based on this analysis, we propose Uncertainty DMD, a simple uncertainty-injection framework that restores stochasticity at two key stages of AR generation: a timestep perturbation for the first chunk to increase first-chunk diversity, and a stochastic cache-writing mechanism for later chunks to preserve uncertainty in autoregressive conditioning. The method requires no architectural changes and introduces only lightweight perturbation operations. The same perturbation mechanisms are used during both training and inference. Experiments show that Uncertainty DMD consistently improves diversity and motion dynamics while maintaining comparable per-sample visual quality.

CommentsProject page: https://scdzx.github.io/Uncertainty-DMD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑