发表机构
Joy Future Academy(京东探索研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自回归视频扩散方法的历史上下文梯度差距问题,提出自梯度强制(SGF)两阶段训练策略,利用未来视频潜变量损失监督模型将上下文编码为有效因果记忆,在长视频外推实验中效果优于自强制。
AI 中文摘要
近期自回归视频扩散方法多基于自强制构建,学生模型在自身展开生成的历史上训练,减少了曝光偏差,但存在历史上下文梯度差距问题。本文提出自梯度强制(SGF),一种两阶段训练策略。第一阶段进行无梯度自回归展开匹配推理并记录相关信息,第二阶段对记录步骤进行并行上下文梯度重建。通过该方法,SGF在原生自回归训练目标中提供缺失的内存写入监督。实验表明,SGF在不同初始化下的长视频外推效果优于自强制,仅用5秒训练窗口就能外推到几分钟的视频。代码和模型将发布以推动自回归视频生成研究。
英文摘要
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.
CommentsProject page: https://zhuang2002.github.io/SelfGradientForcing/