发表机构
State Key Laboratory for Novel Software Technology, Nanjing University; Tsinghua University; Kuaishou Technology; Shanghai Artificial Intelligence Laboratory(南京大学计算机软件新技术国家重点实验室; 清华大学; 快手科技; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流式视频生成的中 horizon 遗忘问题,提出Memory Forcing方法,通过Archive & Working Banks划分缓存,在固定内存下保留历史,在长片段生成中表现优异,且可扩展至Wan2.2 5B模型。
AI 中文摘要
自回归视频扩散支持因果流式视频生成,无需对完整片段进行双向处理,但现有少步系统通常仅在固定大小的KV缓存中保留开头和最新帧。一旦事件离开该窗口,后续帧便无法再关注它,我们将此缺陷称为中 horizon 遗忘。我们提出记忆强制(Memory Forcing),这是一种少步流式方法,可在不增加缓存大小的情况下保留缺失的历史。其存档与工作库(Archive & Working Banks)将缓存划分为汇、存档和工作区域,在固定内存下保留不同的中间事件及近期运动。由于绝对时间索引会偏离训练范围,感知库的旋转位置编码(Bank-aware RoPE)会在注意力计算时重新分配索引,使每个库保持可区分性。在13亿参数规模下,Memory Forcing在更长片段上表现领先,在报告所有四种长度的方法中,其从5秒到60秒的下降幅度最小,且能在对象离开再返回时保留主体和场景。相同设计可扩展至50亿参数的Wan2.2,生成更符合物理规律、更真实且动态的视频,据我们所知,这是该强制路线上首个公开的50亿参数模型。
英文摘要
Autoregressive video diffusion enables causal video streaming without a bidirectional pass over the full clip, but existing few-step systems usually retain only the opening and most recent frames in a fixed-size KV cache. Once an event leaves this window, later frames can no longer attend to it, a failure we term mid-horizon forgetting. We present Memory Forcing, a few-step streaming method that preserves this missing history without increasing the cache size. Its Archive \& Working Banks partition the cache into sink, archive, and working regions, retaining diverse intermediate events alongside recent motion under fixed memory. Because absolute temporal indices drift outside the training range, Bank-aware RoPE reassigns indices at attention time so each bank remains distinguishable. At 1.3B, Memory Forcing leads on longer clips, shows the smallest drop from 5s to 60s among methods reporting all four lengths, and preserves subjects and scenes through leave-and-return. The same design scales to Wan2.2 5B, producing more physically plausible, realistic, and dynamic videos and, to our knowledge, the first public 5B model on this forcing line.
Comments10 pages, 6 figures