arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持久计算状态:用于生成式世界模型的以会话为中心的运行时

Persistent Computational State: A Session-Centric Runtime for Generative World Models

Zhen Lin

arXiv 2607.21686首次发表:更新:

AI 中文总结

研究发现当前视频世界模型用作模拟器失败的归因不完整,定义了持久计算状态(PCS),构建以会话为中心的运行时,通过测量发现PCS,实现低检查点和恢复成本以及新的内存管理方式。

AI 中文摘要

生成式世界模型越来越多地被用作模拟器:规划器分叉一个状态,展开未来,回溯,并返回到访问过的视点。最近的基准测试表明,当前的视频世界模型在这种用法上失败了,并将其归因于模型本身,从而规定了新的架构和训练目标。我们表明这种归因是不完整的,对于一类重要的模型来说甚至是错误的。通过架构对运行时已经持有的状态(一个观察值加上随机数生成器状态、一个内存库或一个窗口化的键值上下文)进行快照,并在真正的偏移之后恢复它,会在所有这三个方面以字节相同的方式重现从未离开的延续;仅损坏随机数生成器会使其退化。这种能力从未缺失:以请求为中心的服务丢弃了它,继承了语言模型服务中运行时状态可重新计算的假设——但世界模型状态携带了一个不可重新计算的内核。我们定义了持久计算状态(PCS),即必须在请求之间存活的最小不可重新计算状态,表明它可以通过测量发现,并在此基础上构建一个以会话为中心的运行时。针对1.85秒的生成步骤,检查点和恢复成本为0.012毫秒;驻留会话成为主机而非设备受限(测量到1024);并且世界内存必须根据与返回的相关性而不是最近性来逐出——这与语言模型的做法相反。

英文摘要

Generative world models are increasingly driven as simulators: a planner forks a state, rolls out futures, backtracks, and returns to a visited viewpoint. Recent benchmarks establish that current video world models fail this usage, and attribute it to the model, prescribing new architectures and training objectives. We show this attribution is incomplete, and for an important class of models simply wrong. Snapshotting the state the runtime already holds -- an observation plus RNG state, a memory bank, or a windowed KV context, by architecture -- and restoring it after a genuine excursion reproduces the never-left continuation byte-identically on all three; corrupting only the RNG degrades it. The capability was never missing: request-centric serving discarded it, inheriting from language-model serving the assumption that runtime state is recomputable -- but world-model state carries a non-recomputable kernel. We define Persistent Computational State (PCS), the minimal non-recomputable state that must survive across requests, show it can be discovered by measurement, and build a session-centric runtime over it. Checkpoint and restore cost 0.012 ms against a 1.85 s generation step; resident sessions become host- rather than device-bounded (measured to 1,024); and world memory must be evicted by relevance to the return, not recency -- the inverse of LLM practice.

Comments29 pages, 8 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑