发表机构
KAIST; Georgia Tech(韩国科学技术院; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程视频生成中动态退化问题,提出LongTake两阶段训练流程,利用长时程教师强制直接初始化DMD,在不增加中间蒸馏步骤下显著提升30秒和60秒滚动生成的动态程度,同时保持美学质量。
AI 中文摘要
世界模型、游戏模拟器以及长镜头视频创作需要连贯的场景演变和长时间内的持续动态。自回归(AR)视频扩散为长时程生成提供了自然框架,然而扩展的滚动生成往往变得近乎静态或损失视觉质量。我们假设这些失败反映了短视频监督对长时间内场景动态如何发展的指导有限。这促使我们引入LongTake,一个围绕在精选真实长视频上进行长时程教师强制(TF)构建的两阶段训练流程。长时程TF训练AR模型基于长真实视频前缀预测后续帧,将直接监督扩展到短训练时程之外。这种监督旨在帮助模型在长时程生成中维持动态并保持视觉质量。我们的核心发现是,该训练阶段增强了在学生自滚动下进行分布匹配蒸馏(DMD)的直接初始化,而无需标准流程中的中间少步蒸馏阶段。在相同的五秒DMD训练设置下,我们的初始化在30秒滚动生成中相比短时程TF初始化,在可比较的美学质量下产生显著更高的动态程度,并在两项指标上均超过评估的基线。混合DMD进一步复用该教师模型,将监督扩展到自滚动的后续帧,同时在初始窗口保留双向联合监督。在长时程自滚动生成中,LongTake位于动态程度与美学质量的帕累托前沿,混合DMD在30秒和60秒时均达到评估方法中最高的动态程度。
英文摘要
World models, game simulators, and long-take video creation require coherent scene evolution and sustained dynamics over extended durations. Autoregressive (AR) video diffusion provides a natural framework for long-horizon generation, yet extended rollouts often become near-static or lose visual quality. We hypothesize that these failures reflect the limited guidance provided by short-video supervision on how ongoing scene dynamics develops over longer durations. This motivates us to introduce LongTake, a two-stage training pipeline built around Long-Horizon Teacher Forcing (TF) on curated real long videos. Long-Horizon TF trains the AR model to predict later frames conditioned on long ground-truth video prefixes, extending direct supervision beyond the short training horizon. This supervision is designed to help the model sustain dynamics and preserve visual quality during long-horizon generation. Our central finding is that this training stage strengthens direct initialization for distribution matching distillation (DMD) under student self-rollout, without the intermediate few-step distillation stage used in standard pipelines. Under the same five-second DMD training setup, our initialization yields substantially higher dynamic degree than short horizon TF initialization on 30-second rollouts at comparable aesthetic quality, and surpasses the evaluated baselines in both measures. Hybrid DMD further reuses this teacher to extend supervision to later frames of the self-rollout while retaining bidirectional joint supervision over the initial window. On long-horizon self-rollouts, LongTake lies on the Pareto front of dynamic degree and aesthetic quality, and Hybrid DMD attains the highest dynamic degree among evaluated methods at both 30s and 60s.