发表机构
Meta Circle; Peking University; Tsinghua University(元环智能; 北京大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示传统世界模型在元胞自动机上失败源于空间/时间局部性和时间稳定性缺失,通过三种信息流修改分别将完成率提升至100%、99.9%和99.9%。
AI 中文摘要
尽管传统世界模型——基于Transformer或卷积网络的自回归或扩散模型——可能学习世界动态的表面统计,但它们能否从观测历史中学习精确的世界动态?利用元胞自动机作为简单测试平台,我们发现许多情况下答案是否定的。传统架构能正确预测大多数像素,但很少能完成完整展开:CNN能预测96.3%的细胞,但仅完成18.9%的展开;联合扩散模型则一个也未完成。我们将这一差距追溯到这些世界模型的三种失败模式——即它们未能精确捕捉空间局部性、时间局部性或时间稳定性。简单的修改即可修复每一种:(1)对于空间局部性,二维旋转位置将Transformer在生命游戏上的完成率从39.1%提升至100%;(2)对于时间局部性,将每个token的细胞前一帧邻域信息赋予该token,将同一Transformer在未见规则上的完成率从25.8%提升至99.9%;(3)对于时间稳定性,因果冻结将同一扩散权重从42.2%提升至99.9%。这三项修改均未触及架构主干,仅改变了其中的信息流。我们还比较了台球上的联合采样和有序采样,并在探索性研究中比较了模拟Burgers方程上的采样。
英文摘要
Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in many cases. Conventional architectures predict most pixels correctly yet rarely complete a rollout: a CNN predicts 96.3% of cells but completes 18.9% of rollouts; a joint diffusion model completes none. We trace the gap to three failure modes of these world models - namely, they fail to exactly capture spatial locality, temporal locality or temporal stability. Simple changes repair each: (1) for spatial locality, two-dimensional rotary positions lift a transformer from 39.1% to 100% on the Game of Life; (2) for temporal locality, handing each token its cell's previous-frame neighbourhood lifts the same transformer from 25.8% to 99.9% on unseen rules; (3) for temporal stability, causal freezing lifts the same diffusion weights from 42.2% to 99.9%. None of the three changes touches the architectural backbone; each only modifies the information flow within it. We also compare joint and ordered sampling on billiards and, in an exploratory study, on a simulated Burgers equation.
Comments35 pages, 18 figures. Code and reproduction materials: https://github.com/guoshaoyang-pku/momentum-induction