LOCI:用于流式世界模型的空间线性记忆
LOCI: Spatial Linear Memory for Streaming World Models
浏览论文内容
中文总结 AI 辅助
LOCI提出混合空间记忆架构,结合键值缓存与基于相机几何的循环线性注意力,在流式世界模型中实现对重访场景的忠实重现,并降低内存开销。
中文摘要 AI 辅助
当相机重新访问先前观察过的区域时,视频世界模型应重现该区域之前的内容。这既需要记住过去的观察结果,也需要为当前视角检索正确的观察结果。键值缓存保留了视觉细节,但会随视频长度增长;循环记忆紧凑,但会将历史压缩为固定大小的状态,因此单个过去的观察结果不再可直接访问。我们引入了LOCI,一种混合空间记忆架构,同时保留这两种表示。在一半的Transformer块中,主注意力保留过去观察的键值缓存;在另一半中,主注意力仅限于当前块,并由循环线性注意力记忆补充,其读写操作以投影相机几何为条件,因此视角同时进入记忆寻址和存储内容。循环读出流入后续的缓存支持的块,并为它们的查询提供累积的场景上下文。在公共MIND记忆基准和保留的录制轨迹上,LOCI比代表性世界模型和同配方全softmax模型更忠实地重现了重新访问的内容;在完整历史下,它在相同长度下将峰值内存相对于全softmax降低了约30%。在有限的保留观察库下,它以恒定内存流式处理长视频,并在相同预算下比全softmax更忠实。
英文摘要
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
发表机构
- Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence(基础模型研究所,穆罕默德·本·扎耶德人工智能大学)
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- Pinscreen
机构由 AI 辅助整理,请以论文原文为准。