发表机构
Alaya Lab; The University of Tokyo(阿莱雅实验室; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究长视频生成中资源分配问题,提出惊喜强制框架,含惊喜门控内存库、基于优先级的替换和相关性感知路由,以及惊喜感知去噪,实验证明该框架提高了长时一致性、视觉质量并保持实时流吞吐量。
AI 中文摘要
流式自回归扩散使分钟级视频合成成为可能,但其有限的上下文和固定的去噪时间表在高度非平稳序列中均匀分配资源。滚动键值缓存会遗忘遥远的视觉证据,且每个生成块不论实际难度都接受相同次数的去噪。我们引入惊喜强制,一个无训练框架,将这些限制视为在线资源分配问题。惊喜门控内存库用值令牌描述符总结被逐出帧,用互补全局偏差和最近邻新奇信号评估,通过归一化分数空间中的反馈控制预算调节准入。基于优先级的替换和相关性感知路由保持外部内存紧凑有用。同时,惊喜感知去噪根据首次去噪后的最大相邻帧余弦距离估计块难度,用局部百分位数调度器跳过较简单块的中间步骤。在VBench、VBench-Long和VBench-2.0上的实验表明,该策略提高了长时一致性和视觉质量,同时保持实时流吞吐量。
英文摘要
Streaming autoregressive diffusion makes minute-scale video synthesis practical, but its bounded context and fixed denoising schedule allocate resources uniformly across a highly non-stationary sequence. A rolling key-value cache forgets distant visual evidence even when that evidence remains important, while every generated chunk receives the same number of denoising passes irrespective of its actual difficulty. We introduce Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems. A Surprise-Gated Memory Bank summarizes evicted frames with value-token descriptors, evaluates them using complementary global-deviation and nearest-neighbor novelty signals, and regulates admission through a feedback-controlled budget in normalized score space. Priority-based replacement and relevance-aware routing then keep the external memory compact and useful. In parallel, Surprise-Aware Denoising estimates chunk difficulty from the maximum adjacent-frame cosine distance after the first denoising pass and uses a local percentile scheduler to skip intermediate steps for comparatively easy chunks. Experiments on VBench, VBench-Long, and VBench-2.0 show that the proposed allocation strategy improves long-horizon consistency and visual quality while retaining real-time streaming throughput.
CommentsTechnical report