Memorizon:训练超越上下文窗口的世界模型
Memorizon: Training World Models Beyond Their Context Window
- IFM, MBZUAI(IFM,穆罕默德·本·扎耶德人工智能大学)
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Memorizon通过共享前向传播和检索式记忆库,解耦监督跨度与注意力成本,使长跨度训练高效,显著提升流式世界模型的重访一致性。
AI中文摘要:
流式世界模型应在重复访问时一致地渲染一个地点。直接监督这种重访需要捕获两次访问的训练样本,这些样本通常跨越数分钟。然而,对完整时间跨度的密集注意力会产生二次方成本,使得长跨度监督变得昂贵。Memorizon打破了这种耦合:长跨度是监督所需的,但并非注意力所需,因为两次访问可以共享一次前向传播,而无需包含每个中间帧。一个训练样本覆盖任意长度的跨度,但仅对其最后$k$个块进行评分。每个被评分的块不是对其之前的历史进行分词,而是通过相机共视性检索自己的前$K$个潜在变量,这些请求的并集形成一个共享库。该库以$kK$为界,因此无论跨度多长,序列都保持有界;在最短跨度下,该配方与常规训练完全相同。添加库会使一步的成本增加一次;除此之外,更长的跨度成本很小,从100秒增加到400秒仅使步时增加12%。与滑动窗口基线相比,检索在每个分割上都提高了重访一致性,而一个足够长以到达每次返回的首次访问的跨度进一步增加了24%至30%,但以一定的图像质量为代价;超过该跨度,更多长度不再有帮助。从另一集填充库会使重访相关性降低83%,因此模型会使用其检索到的内容。项目页面:此https URL
英文摘要:
Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last $k$ chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-$K$ latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by $kK$, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon