arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视频世界模型的可寻址内存

Addressable Memory for Video World Models

Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljoša Ošep

arXiv 2608.07408首次发表:更新:

发表机构

NVIDIA; Princeton University; University of Toronto; Vector Institute(英伟达; 普林斯顿大学; 多伦多大学; 矢量研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对交互式视频世界模型超出训练时序后内存寻址失效及压缩缓存破坏内存的问题,提出无训练框架 WorldTrace,含两种压缩方法,在新基准 LoopBench 上分别提升时序一致性 15.5%、情景回忆 19.5%。

AI 中文摘要

我们研究交互式视频世界模型中的视觉持续性。这些模型依赖键值(KV)缓存作为不断增长的视觉内存来承载之前生成的帧。然而,我们发现当 rollout(展开)超出训练 horizon(时序范围)时,模型无法可靠地寻址存储的内容,因为时序旋转位置嵌入(RoPE)偏移会超出训练期间见过的范围,模型难以通过注意力机制检索相关视觉信息。此外,在 RoPE 旋转空间中朴素地压缩缓存会通过平均不兼容的位置相位破坏内存。为解决该问题,我们提出 WorldTrace,一种用于长时序视觉持续性的无训练内存框架。WorldTrace 通过为每个摘要槽分配一个不同的、符合分布的虚拟位置来保持压缩内存的可寻址性。在该可寻址缓存内,我们研究两种内存压缩方法:WorldTrace-Field 压缩历史以实现时序一致性,WorldTrace-Landmark 在检测到的转换处存储逐字场景轨迹以实现 episodic recall(情景回忆)。我们进一步引入 LoopBench,一个评估压缩缓存能否在长绕行后重建先前访问场景的基准。在 LoopBench 上,WorldTrace-Field 将时序一致性提升了 +15.5%,WorldTrace-Landmark 将情景回忆提升了 +19.5%,无需重新训练即可扩展视觉持续性生成。

英文摘要

We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.

CommentsProject page: https://research.nvidia.com/labs/sil/projects/WorldTrace/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑