arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02339cs.LGcs.AI

AGI迷宫预测数据集:用于学习Transformer世界动力学的紧凑基准

AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers

  • SingularityNET Foundation(奇点网基金会)

机构由 AI 辅助整理,请以论文原文为准。

Alexey Potapov

中文总结 AI 辅助

该研究推出AGI迷宫预测数据集与基准,通过构建伪视频空间记忆Transformer模型,证实结构化工作记忆比单纯增加隐式容量更有效,为测试文本接口与结构化状态耦合的架构提供了紧凑场景。

中文摘要 AI 辅助

世界建模需要预测模型维持并更新内部状态,以充分推理行动的后果。我们推出AGI迷宫预测数据集与基准,这是一个轻量受控测试平台,用于研究Transformer及其他预测模型的该能力。该基准源自程序生成的有状态网格世界,包含每步转移预测、固定视界状态预测及序列文本观测预测。源迷宫不相交的训练与验证划分,结合贪心精确匹配评估,可区分学习可迁移的行动条件动力学与记忆熟悉布局中的转移。我们建立了从零开始的字节级Transformer基线,并将其与两种增强工作记忆的架构对比。通用的辅助隐式记忆Transformer可完美拟合部分训练集,但无法持续提升保留的性能。相比之下,伪视频空间记忆Transformer从输入地图初始化二维隐式工作区,并根据行动历史更新,无需接收中间地图、位置或状态标签。在相同数据、目标与评估协议下,该模型在选定的固定视界任务上达到完美验证准确率,而字节级与非结构化记忆基线则未做到,且大幅提升了序列文本轨迹预测性能。这些结果表明,结构化、与任务对齐的工作记忆可能比单纯增加隐式容量更有用。更广泛地说,我们认为语言接地由持久数据结构及对其的计算介导;该基准提供了一个紧凑场景,用于测试将文本接口与学习到的结构化状态耦合的架构。

英文摘要

World modeling requires a predictive model to maintain and update an internal state adequate for reasoning about the consequences of actions. We introduce the AGI Maze Prediction Datasets and Benchmark, a lightweight controlled testbed for studying this capability in Transformers and other predictive models. Derived from procedurally generated, stateful grid worlds, the benchmark comprises per-step transition prediction, fixed-horizon state prediction, and sequential textual-observation prediction. Source-maze-disjoint training and validation splits, together with greedy exact-match evaluation, distinguish learning transferable action-conditioned dynamics from memorizing transitions in familiar layouts. We establish from-scratch byte-level Transformer baselines and compare them with two working-memory-augmented architectures. A generic auxiliary latent-memory Transformer can fit some training sets perfectly but does not consistently improve held-out performance. In contrast, a pseudo-video spatial-memory Transformer initializes a two-dimensional latent workspace from the input map and updates it from action history without receiving intermediate maps, positions, or state labels. Under the same data, objectives, and evaluation protocol, this model reaches perfect validation accuracy on selected fixed-horizon tasks where the byte and unstructured-memory baselines do not, and substantially improves sequential text-trace prediction. These results suggest that structured, task-aligned working memory can be more useful than additional latent capacity alone. More broadly, we argue that language grounding is mediated by persistent data structures and computations over them; the benchmark offers a compact setting for testing architectures that couple textual interfaces to learned structured state.

↑