AI 中文总结
本文针对自回归长视频生成中的误差累积问题,提出无需训练的FreqForcing框架,通过频谱自锚定技术实现24倍外推,性能优于现有无训练方法且与有训练方法有竞争力。
AI 中文摘要
自回归视频扩散模型支持实时流式视频生成,但自回归过程中产生的误差会在长时序范围内累积,表现为颜色漂移、运动停滞,最终导致视觉崩溃。本文从频域角度分析该现象:误差累积体现为低频带的明显能量漂移。进一步研究频域中的注意力汇(attention sink)的有效性,发现其虽能在一定程度上缓解频谱能量漂移以提升视频质量,但无法完全解决问题。基于上述分析,本文提出FreqForcing,这是一种无需训练的框架,通过频谱自锚定(Spectral Self-Anchoring,SSA)解决长视频生成中的误差累积问题。所提SSA利用锚点注意力的低频分量维持长时序视觉稳定性,同时通过局部注意力的高频分量保留动态运动。本文的FreqForcing将在5秒片段上预训练的自强迫(Self-Forcing)方法扩展至2分钟视频生成,实现了24倍外推。大量实验表明,FreqForcing在定量和定性上均优于现有无需训练的方法,同时与代表性的基于训练的方法具有竞争力。
英文摘要
Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long horizons, manifesting as color drift, motion stagnation, and eventual visual collapse. In this paper, we characterize this phenomenon from a frequency-domain perspective: error accumulation appears as a pronounced energy drift in the low-frequency bands. We further investigate the effectiveness of attention sink in the frequency domain, and find that it improves the video quality by alleviating the spectral energy drift to some extent, but cannot fully resolve it. Motivated by the above analysis, we propose FreqForcing, a training-free framework that addresses error accumulation in long-video generation via Spectral Self-Anchoring (SSA). The proposed SSA leverages the low-frequency components of anchor attention to maintain long-horizon visual stability, while preserving dynamic motion through the high-frequency components of local attention. Our FreqForcing extends Self-Forcing pretrained on 5s clips to two-minute generation, achieving 24x extrapolation. Extensive experiments show that FreqForcing outperforms existing training-free methods quantitatively and qualitatively while remaining competitive with representative training-based approaches.
CommentsCode is available at: https://github.com/jiatongli2024/FreqForcing