arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38114cs.CV

自对齐强制:基于可微噪声历史的流式视频扩散

Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

Weiqiang Wang, Zhuokun Chen, Yusheng Dai, Boying Li, Yi Zhang, Hossein Rahmani, Qiuhong Ke, Jianfei Cai

首次发表
浏览论文内容

中文总结 AI 辅助

提出自对齐强制(SAF)训练方案,通过将历史噪声水平与去噪块对齐,实现可微历史,加速训练并提升长时程视频生成质量。

中文摘要 AI 辅助

自回归视频扩散能够实现交互式流式生成,但在长时间展开过程中存在误差累积问题。自滚动训练减少了曝光偏差,然而有限的滚动长度仍无法解决长距离漂移问题。我们观察到,历史键值(K/V)表示的噪声水平在视觉质量与运动之间进行权衡,并且通过历史恢复梯度使得因果训练与双向训练的契合度大大提高。基于这些观察,我们引入了自对齐强制(SAF),一种训练方案,将每个块的历史与其去噪块的噪声水平对齐。具体而言,历史是由同一去噪阶段的前置块产生的K/V,因此一个阶段的所有块可以在因果掩码下通过单次前向传播完成去噪。这保持了噪声历史的可微性,使得未来的损失能够优化其编码方式。因此,SAF避免了单独的无梯度滚动和每块时间步零的重新缓存,训练速度比先前方法快达1.8倍,且内存占用更低。在推理时,SAF在现有方法中实现了最高的单GPU吞吐量,并为多GPU流水线保留每阶段一个历史库,在4个GPU上达到49.1 FPS。实验表明,SAF在长时程生成方面表现优越,并在视觉质量与运动之间取得了更好的平衡。项目页面:此https URL。

英文摘要

Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.

发表机构

  • Monash University(莫纳什大学)
  • Vivix AI
  • Lancaster University(兰卡斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑