发表机构
Northeastern University; Adobe Research(东北大学; Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ViRDM,一种无需教师和评论模型的视频后训练方法,通过表示分布匹配和轻量级正则化,在少步因果视频生成中实现高质量和低资源消耗,仅20次更新即在VBench上达到84.87分。
AI 中文摘要
少步自回归(AR)视频扩散模型能够实现低延迟流式生成,但现有的后训练方法主要依赖分布匹配蒸馏(DMD),这需要大型预训练教师模型和在线评论模型,通过扩散分数来估计分布差异。在这项工作中,我们探究是否可以通过仅针对预计算的目标分布对生成器进行后训练,来消除这种资源密集型的教师-评论模型堆栈。受一步图像生成中表示分布匹配(RDM)的启发,我们系统地研究了其向少步因果视频生成的迁移,并识别出三个关键障碍:内存不可行的梯度路径、独特的视频优化机制,以及对时间动态约束不足的表示分布。我们提出了ViRDM,一种无需教师和评论模型的视频后训练方法,该方法依次解决这些障碍。通过将RDM与随机截断的干净出口监督、轻量级VAE解码器以及分阶段向量-雅可比乘积相结合,ViRDM使表示分布匹配在多步因果视频展开中内存可行。我们进一步建立了视频RDM的有效生成群体和初始化机制,并引入轻量级动态正则化以弥补时间动态约束不足。ViRDM将三网络蒸馏转变为仅生成器后训练,减少了GPU内存使用和训练时间,同时提高了视频质量。仅需20次生成器更新,该方法在官方VBench评估中达到84.87分,比之前最好的少步因果基线高出0.36分,同时仅需16个A100 GPU小时。我们还报告了探索性结果,展示了相同方法在更低因果采样预算以及一步、两步和四步双向生成中的潜力。
英文摘要
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained teacher and an online critic to estimate distributional discrepancies through diffusion scores. In this work, we ask whether this resource-intensive teacher--critic stack can be eliminated by post-training only the generator against a precomputed target distribution. Drawing inspiration from representation distribution matching (RDM) for one-step image generation, we systematically study its transfer to few-step causal video generation and identify three key barriers: a memory-intractable gradient path, a distinct video optimization regime, and representation distributions that underconstrain temporal dynamics. We introduce ViRDM, a teacher- and critic-free video post-training recipe that addresses these barriers sequentially. By coupling RDM with stochastically truncated clean-exit supervision, a lightweight VAE decoder, and staged vector--Jacobian products, ViRDM makes representation distribution matching memory-feasible for multi-step causal video rollouts. We further establish effective generated-population and initialization regimes for video RDM, and introduce lightweight dynamics regularization to compensate for the underconstrained temporal dynamics. ViRDM turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality. With only 20 generator updates, the recipe reaches 84.87 on the official VBench evaluation, outperforming the previous best few-step causal baseline by 0.36, while requiring 16 A100 GPU-hours. We additionally report exploratory results demonstrating the potential of the same recipe for lower causal sampling budget and for one-, two-, and four-step bidirectional generation.
CommentsTech Report