从分数到样本:自回归视频生成的弹性强制
From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
浏览论文内容
中文总结 AI 辅助
本文提出弹性强制方法,直接从参考视频学习自回归生成分布,消除蒸馏中的分数模型,提升VBench总分并支持更大规模训练。
中文摘要 AI 辅助
少步自回归视频生成通常依赖于分布匹配蒸馏(DMD),这需要双向扩散教师模型和在线假分数模型。我们则直接从参考视频中学习 rollout 分布,在训练后阶段消除了这两个分数模型。我们的框架在冻结的自监督视频表示空间中最小化最大均值差异(MMD),使用混合 Nyström--Monte Carlo 估计器来平衡近似偏差与采样方差。内存高效的重放和梯度子采样使这一目标变得实用。使用与 Self-Forcing 相同的架构和初始化,我们的 1.3B 模型将 VBench 总分从 83.80 提升到 84.64,同时保持 17 FPS。移除辅助分数模型还使得在八块 H200 GPU 上进行 14B 模型的训练后优化成为可能。除了蒸馏之外,从参考视频中学习使得无需特定目标的扩散教师模型即可获取新的视觉风格、语义概念和空间先验。
英文摘要
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
发表机构
- College of AI, Tsinghua University(清华大学人工智能学院)
- IAIR, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所)
- Xianghui Academy, Fudan University(复旦大学相辉学院)
- Peking University(北京大学)
- BAAI(北京智源人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。