arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10429cs.CV

SGF+:自回归视频生成中的梯度流解耦

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

  • Tsinghua University(清华大学)
  • Joy Future Academy(京东探索研究院)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Zihan Su, Junhao Zhuang, Yaowei Li, Siwen Lu, Haoran Li, Lingen Li, Haoyu Wu, Weiyang Jin, Songchun Zhang, Haoyang Huang, Chun Yuan, Zeyue Xue, Nan Duan

中文总结 AI 辅助

SGF+通过为自回归视频生成中的上下文写入和去噪分配独立参数,解耦梯度流,在不增加训练数据的情况下提升视觉质量与长时程一致性,并支持长达24小时的连续生成。

中文摘要 AI 辅助

自回归视频生成需要对当前帧进行去噪,同时将其键值表示写入上下文以供未来预测使用。然而,这两个角色通常共享参数,我们发现它们的梯度表现出不同的模式和系统性负对齐,阻碍了视觉质量和时间一致性的联合优化。我们引入了自梯度强制加(SGF+),该方法为上下文写入和去噪分配独立的参数,同时通过因果注意力保持它们之间的交互。两个角色都使用原始生成目标进行联合优化,无需辅助损失,上下文写入通过其对未来预测的贡献进行监督。这一简单改变在逐帧和分块生成中,相较于评估的基线方法,提升了视觉质量和长时程一致性,且无需额外的视频训练数据或更长的训练周期。仅使用5秒的片段进行训练,SGF+支持长达24小时的连续生成,而无需长视频微调。这些结果凸显了角色特定参数化作为高质量自回归视频生成和原生长时程外推的有效设计原则。

英文摘要

Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.

补充信息

↑