arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习检索:内化视频世界模型的记忆检索

Learning to Retrieve: Internalizing Memory Retrieval for Video World Models

JiaKui Hu, Tailai Chen, Yuqi Pan, Xuerui Qiu, Jialun Liu, Xiao Cao, Zhenxin Zhu, Guang Chen, Hangjun Ye, Bing Wang, Yanye Lu

arXiv 2610.11444首次发表:更新:

发表机构

Peking University; Xiaomi EV; CASIA; NUS(北京大学; 小米汽车; 中国科学院自动化研究所; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Learning-to-Retrieve(L2R)方法,将记忆检索内化到视频世界模型的生成过程中,无需外部记忆系统,即可提升长序列生成的场景一致性。

AI 中文摘要

视频世界模型旨在生成由相机轨迹条件控制的可探索、3D一致的场景视频。现有方法通常依赖外部记忆系统,在长序列生成过程中显式检索先前观测到的内容以缓解场景漂移。然而,这些辅助记忆通路在模型的内部生成动态之外运行,导致模型无法从本质上学习何时以及应检索何种历史信息。我们提出将记忆检索内化到生成过程中,使检索作为视频世界模型的内在行为出现,而非依赖外部记忆系统。基于此,我们引入学习检索(Learning-to-Retrieve, L2R),将模型的持久内部状态重新用作历史上下文的记忆。相机条件检索门从该状态中选择性访问相关历史信息,确定“要检索什么”,而检索触发器确定“何时检索”。我们进一步用3D重可见性信号监督触发器,当先前观测到的内容重新进入当前视图时激活检索,否则保留现有上下文。这些组件共同使模型从本质上获得记忆检索行为,并将相关历史观测纳入生成过程,无需单独的检索通路。在多个基础模型和相机重访基准上,L2R提升了长期场景一致性,同时消除了对外部记忆库或3D条件的需求。

英文摘要

Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories. Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation. However, these auxiliary memory pathways operate outside the model's internal generative dynamics, preventing the model from intrinsically learning when and what historical information should be retrieved. We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system. Based on this principle, we introduce \textbf{Learning-to-Retrieve (L2R)}, which repurposes the model's persistent internal state as a memory for historical context. A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining \textit{what to retrieve}, while a retrieval trigger determines \textit{when to retrieve}. We further supervise the trigger with a 3D re-visibility signal, activating retrieval when previously observed content re-enters the current view while otherwise preserving the existing context. Together, these components enable the model to intrinsically acquire memory retrieval behavior and incorporate relevant historical observations into generation without a separate retrieval pathway. Across multiple base models and camera-revisit benchmarks, L2R improves long-term scene consistency while eliminating the need for an external memory bank or 3D conditions. https://jkhu29.github.io/l2r

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑