arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32540cs.CVcs.AIcs.LGcs.MM

飞行中KV缓存与干净锚点:加速自回归视频扩散

In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang

首次发表
浏览论文内容

中文总结 AI 辅助

FlashForward通过重用去噪阶段的飞行中KV缓存并辅以稀疏干净锚点,加速自回归视频扩散,实现多GPU并行,提升速度与质量。

中文摘要 AI 辅助

少步自回归视频扩散通过将长视频分割为时间块并逐块生成,每个块经过短序列的去噪阶段,从而生成长视频。为了记忆已生成的块,先前的方法通过额外的前向传播重建干净或较少噪声的键值(KV)缓存,以在不推进输出潜在变量的情况下构建缓存。然而,每次去噪前向本身已经计算了当前块的飞行中KV。我们提出FlashForward,直接重用此缓存以避免仅用于缓存更新的重型模型前向。在当前块完成一个去噪阶段后,其阶段特定缓存即可用于下一个块。因此,将每个阶段分配给一个GPU,可以让不同块同时占用不同阶段。这种早期可用性带来质量代价:由此产生的阶段匹配历史是噪声的,导致块间外观和运动漂移。为弥补此不足,FlashForward在相应区域生成之前产生稀疏的辅助干净锚点潜在变量,从而通过双侧条件稳定生成轨迹。两种记忆在不同时间尺度上运行:稀疏干净锚点KV提供粗粒度、长距离的双侧结构指导,而密集阶段匹配历史保留细粒度、近期演化。使用最多四个GPU,FlashForward在1.3B和14B骨干规模、480p和720p分辨率下,对于20秒或更长的16 FPS视频,运行速度比HiAR快1.16--1.69倍,比Self-Forcing快1.42--2.92倍。在VBench上,对于480p的1.3B模型,它获得更高分数并在更长时长下保持稳定,表明FlashForward在20秒、35秒和65秒时长下以更快的生成速度生成高质量且时间一致性的视频。

英文摘要

Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. Assigning one GPU to each stage therefore lets different chunks occupy different stages concurrently. This early availability has a quality cost: the resulting stage-matched history is noisy, causing appearance and motion drift among chunks. To complement it, FlashForward produces sparse auxiliary clean anchor latents before the corresponding region is generated so the generation trajectories can be stabilized by this two-sided conditioning. The two memories operate at different temporal scales: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance, while dense stage-matched history preserves fine, recent evolution. With up to four GPUs, FlashForward runs $1.16$--$1.69\times$ faster than HiAR and $1.42$--$2.92\times$ faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbone scales at 480p and 720p. On VBench, for the 1.3B model at 480p, it achieves higher scores and remains stable at longer durations, demonstrating that FlashForward generates high-quality and temporally consistent videos across durations of 20s, 35s and 65s at a much faster generation speed.

补充信息

↑