MosaiChunk:为自回归视频生成组合时空记忆
MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation
浏览论文内容
中文总结 AI 辅助
MosaiChunk通过组合历史键值条目的时空记忆,在固定预算下提升长时程自回归视频生成的重访一致性,并引入RememBench基准验证其效果。
中文摘要 AI 辅助
长时程自回归视频生成受限于有限的上下文窗口。当某个物体或场景超出上下文范围时,其细粒度的视觉细节可能会丢失,并且在重新出现时难以恢复。为了保留对这些视觉细节的访问,我们引入了MosaiChunk,一种时空记忆机制,它跨空间和时间组合一个由选定的历史键值(KV)条目组成的马赛克。我们的方法基于一个观察:一个冻结的视频生成器可以直接消费这种非连续的历史KV并恢复相应的视觉内容。因此,我们保持生成器固定,仅学习一个轻量级路由器,该路由器在固定的活动记忆预算下决定将哪些历史片段包含在马赛克中。我们进一步引入了RememBench,一个长时程重访基准,包含基于提示的文本到视频(T2V)和基于相机的图像到视频(I2V)两个子集。我们的实验表明,在匹配的记忆预算下,无论是在T2V还是I2V设置中,MosaiChunk在重访一致性上都持续优于滑动窗口推理和整块检索。
英文摘要
Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain access to such visual details, we introduce MosaiChunk, a spatio-temporal memory mechanism that composes a mosaic of selected historical key-value (KV) entries across space and time. Our approach is motivated by the observation that a frozen video generator can directly consume such non-contiguous historical KV and recover the corresponding visual content. We therefore keep the generator fixed and learn only a lightweight router that determines which historical sections to include in the mosaic under a fixed active-memory budget. We further introduce RememBench, a benchmark of long-horizon revisits with prompt-driven text-to-video (T2V) and camera-driven image-to-video (I2V) splits. Our experiments show that MosaiChunk consistently improves revisit consistency over both sliding-window inference and whole-chunk retrieval under matched memory budgets, across both T2V and I2V settings.
发表机构
- Cornell University(康奈尔大学)
- University of California, Berkeley(加州大学伯克利分校)
- Harvard University(哈佛大学)
- Impossible, Inc.(Impossible 公司)
机构由 AI 辅助整理,请以论文原文为准。