arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07077cs.AIcs.LG

Transformer难以利用其涌现的世界模型:重新审视汉诺塔与思维的错觉

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking

Devin Pereira, Willem Zuidema

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究发现,Qwen3.6-27B等大型推理模型虽能编码汉诺塔的谢尔宾斯基世界模型,但因表示衰减导致圆环数超3时失败,注入提示词时间表示可部分恢复性能。

中文摘要 AI 辅助

汉诺塔是一个简单的规划谜题,已有研究表明大型推理模型(LRM)难以解决该问题。当前模型可解决该谜题的标准形式,但仍难以应对平-平变体(初始和目标状态不要求所有圆环都在同一 peg 上)。本文对小型自研 Transformer 和大型第三方 LRM 如何解决该任务展开深入研究。为从机制上理解失败原因,我们首先基于预计算的解决方案轨迹从头训练小型 Transformer,运用多种可解释性技术发现这些 Transformer 会形成涌现的世界模型:即对谜题状态空间(谢尔宾斯基三角形)的线性可解码、几何上忠实的表示,且该表示与谜题求解存在因果关联。其次,我们转向大型语言模型(LLM),将上述技术应用于两个前沿推理模型 Qwen3.6-27B 和 DeepSeek-R1-Distill-Qwen-32B,这两个模型尝试通过扩展思维链来解决该任务。令人惊讶的是,我们发现这两个模型在提示词末尾近乎完美地编码了谢尔宾斯基世界模型,但当圆环数超过3个时,多数任务仍会失败。我们将失败根源定位为世界模型表示的衰减:我们在规划的不同阶段探测该表示,并通过在推理时注入提示词时间的表示来建立因果关系,结果表明性能可得到提升。因此,模型的失败是所需表示的维持问题,而非表示缺失问题,且性能至少部分可恢复。这些结果重新解释了已有研究中报告的性能崩溃:当前大型推理模型会构建世界模型,随后失去它。

英文摘要

The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the standard formulation of the puzzle, but still struggle with the flat-to-flat variant (where initial and goal states are not restricted to have all rings on a single peg). This paper presents an in-depth study of how both small, in-house Transformers and large, third-party LRMs solve this task. To understand the failures mechanistically, we first train small Transformers from scratch on precomputed solution traces. Using a variety of interpretability techniques, we show that these Transformers develop an emergent world model: a linearly decodable, geometrically faithful representation of the puzzle's state space (the Sierpinski triangle), that is causally involved in solving the puzzles. Second, we return to the large LLMs and apply our techniques to two frontier reasoning models, Qwen3.6-27B and DeepSeek-R1-Distill-Qwen-32B, that attempt to solve the task through extended chain-of-thought. Surprisingly, we find that both models encode the Sierpinski world model near-perfectly at the end of the prompt, and yet fail at the majority of tasks when there are more than 3 rings. We locate the source of this failure in the decaying representation of the world model. We probe for the representation at different stages during planning, and establish causality by showing that performance can be improved by injecting the prompt-time representation at inference. The failure of the models is thus one of maintenance of the required representations, not their absence, and performance is at least partially recoverable. These results thus reframe the reported collapse in performance from prior work: current Large Reasoning Models build a world model, and then lose it.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

↑