发表机构
EPFL; IDIAP Research Institute; MATS; University of Bern; University of Manchester(洛桑联邦理工学院; IDIAP研究所; MATS; 伯尔尼大学; 曼彻斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过机制分析揭示Transformer即使行为失败也可能具备世界模型,提出能力打包缓解特征干扰,并倡导机制性研究世界建模能力。
AI 中文摘要
行为失败可能使一个Transformer看起来缺乏世界模型,即使它已经学习了其环境的忠实表征。我们在TaxiGPT中证明了这一点,这是一个在曼哈顿随机游走数据上训练的Transformer,其失败曾被解释为内部地图不连贯的证据。通过机制分析和因果干预,我们表明该模型表征了交叉路口和街道,跟踪其位置,并使用目标罗盘进行导航。我们将其失败追溯到叠加交叉路口特征之间的干扰,这种干扰扰乱了内部地图中的定位。能力打包(Affordance packing),即将具有相同合法移动的交叉路口表征分组,有助于限制这些错误的后果。最后,我们提出了机制指标,用于比较模型,并表明世界建模能力在训练的不同阶段出现。我们的发现促使从询问模型是否具有世界模型转向机制性地研究其世界建模:即模型表征其环境并使用这些表征指导行为的相互作用的能力。
英文摘要
Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.