发表机构
Seoul National University; Roblox(首尔大学; 罗布乐思公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过动作条件视频“贝壳游戏”实验,发现多数视频世界模型无法追踪未观测的世界状态,仅线性注意力(转移特征值为负)和TTT(带非线性快速权重)可实现外推,为构建有状态视频世界模型提供了方向。
AI 中文摘要
视频世界模型正越来越多地被用作模拟器,但仅靠视觉保真度并不能表明模型维持了世界的隐藏状态。我们通过一个动作条件视频“贝壳游戏”来研究这一差距,这是S₅状态追踪的视觉类似物,它将视觉渲染与底层隐藏状态的合成分离开来。双向和自回归Transformer、Mamba以及限制为非负转移特征值的线性注意力,在5次交换的训练范围内均表现良好,但在更长的交换链(外推)上则会降至随机水平,同时仍能渲染出合理的视频,额外的去噪步骤并无益处。基于像素的扩散目标从未对未被观测的隐藏状态进行监督,因此生成的帧无法承载该状态,且该状态必须存在于架构内部而非token中。对于Transformer而言,其架构状态仅是一个仅追加的KV缓存,因此模型必须在每个块中从整个历史中重新推导隐藏排列。我们发现两种可实现外推的机制,两者均在块间传递状态并就地修正该状态:线性注意力在其转移特征值可为负时成功,而带有非线性快速权重的TTT则通过更新其读取自身状态的特征映射来实现。我们进一步在动态世界探索任务中研究了更困难的案例,并讨论了构建有状态视频世界模型的更广泛意义。
英文摘要
Video world models are increasingly used as simulators, but visual fidelity alone does not show that a model maintains the hidden state of the world. We examine this difference with an action-conditioned video Shell Game, a visual analogue of $S_5$ state tracking that separates visual rendering from compositing the unobserved world state. Trained on 5-swap chains, standard backbones (e.g., bidirectional and autoregressive Transformers, Mamba, and linear attention) render plausible videos and predict the correct ball location up to 5 swaps. However, they fail to learn the rule and generalize to longer swap chains, even with more denoising steps. As the pixel-based diffusion loss does not force the generated frames to hold the unseen ball position, output tokens cannot carry it, and the state has to live within the architecture. In a causal Transformer, this implicit state is an append-only KV cache, which is written once and never revised, so the model must re-compose the swaps at every chunk. Tracking $S_5$ this way requires depth to grow with sequence length, which no fixed-depth Transformer provides. We study what enables learning the rule, and find that length generalization requires a revisable state carried across chunks and an update expressive enough to apply a swap. Linear attention can achieve this by allowing negative transition eigenvalues, and autoregressive Transformers can do so with nonlinear TTT fast weights (e.g., SwiGLU) whose online updates change the feature map used to read their state. We further examine Memory Maze and Block World, where the state is not fixed by the input action stream alone and must be corrected from observations or keeps changing out of view, and discuss the implications for building stateful video world models.
CommentsProject webpage:https://joonghyuk.com/stateful-vwm-web/