arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39132cs.CVcs.MM

不确定性感知一致性蒸馏用于少步视频生成

Uncertainty-Aware Consistency Distillation for Few-Step Video Generation

Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

针对少步视频生成中教师目标不可靠的问题,提出不确定性感知一致性蒸馏,用无参数不确定性估计重加权监督,结合特征对抗训练,在VBench 2.0上实现最先进的4步生成。

中文摘要 AI 辅助

我们研究少步视频生成,即将通常需要数十步采样、产生大量延迟和计算的多步视频生成器蒸馏为少步学生模型。一致性蒸馏是一种常用方法,其中多步教师为少步学生提供一致性目标。然而,这些教师引导的目标并非同样可靠,且内容在时间上快速变化的地方(例如移动的树叶阴影或流动的水)更难学习。我们观察到,监督可靠性遵循内容的局部难度而非语义复杂度:变化小的区域产生一致的端点预测,而时间变化大的区域产生更大的差异,这些差异与最大的感知误差相吻合。受此观察启发,我们提出不确定性感知一致性蒸馏(UACD),该方法使用局部、无参数的不确定性估计对每个时空区域的一致性监督进行重新加权。具体而言,我们构建两条独立扰动的教师引导一致性路径,其学生端点预测提供共识目标;学生直接预测与该目标之间的差异作为不确定性代理。然后,我们通过指数权重放宽高不确定性区域的一致性惩罚,同时在其他区域保持完整惩罚,因为学生无法期望匹配难以学习的目标。为了在激进步数减少下保持感知质量,我们集成了特征空间对抗训练与语义对齐。通过对50步Wan模型进行参数高效的LoRA适配,我们的方法在VBench 2.0上实现了最先进的4步生成(平均得分0.556),并在用户研究中优于竞争方法。

英文摘要

We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.

发表机构

  • Hefei University of Technology(合肥工业大学)
  • University of Macau(澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

↑