arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自梯度强制:原生长视频外推

Self Gradient Forcing: Native Long Video Extrapolation

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian, Yaowei Li, Yawen Luo, Yijun Liu, Weiyang Jin, Songchun Zhang, Xianglong He, Xuying Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan

arXiv 2607.20368首次发表:更新:

发表机构

Joy Future Academy(京东探索研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自回归视频扩散方法的历史上下文梯度差距问题,提出自梯度强制(SGF)两阶段训练策略,利用未来视频潜变量损失监督模型将上下文编码为有效因果记忆,在长视频外推实验中效果优于自强制。

AI 中文摘要

近期自回归视频扩散方法多基于自强制构建,学生模型在自身展开生成的历史上训练,减少了曝光偏差,但存在历史上下文梯度差距问题。本文提出自梯度强制(SGF),一种两阶段训练策略。第一阶段进行无梯度自回归展开匹配推理并记录相关信息,第二阶段对记录步骤进行并行上下文梯度重建。通过该方法,SGF在原生自回归训练目标中提供缺失的内存写入监督。实验表明,SGF在不同初始化下的长视频外推效果优于自强制,仅用5秒训练窗口就能外推到几分钟的视频。代码和模型将发布以推动自回归视频生成研究。

英文摘要

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.

CommentsProject page: https://zhuang2002.github.io/SelfGradientForcing/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑