发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过四阶段分解揭示大语言模型数学推理的机制,定位干扰导致失败的关键阶段为操作规划,并双向验证了注意力头的因果作用。
AI 中文摘要
大语言模型能以高准确率解答小学数学应用题,然而在问题中插入一个无关子句即可使其崩溃。我们通过机制性解释调和了这些观察。我们表明,模型的内部计算可分解为一个四阶段顺序流水线:模式抽象、操作规划、操作数绑定与计算,每个阶段在可识别的层带中产生不同的中间表示。利用同一框架来诊断干扰引起的失败,我们将破坏定位到单一阶段——操作规划,该阶段由一组注意力头实现,我们通过双向验证了其因果作用。简言之,我们为大语言模型中的数学应用题推理提供了机制性解释,并解释了其受干扰时的失败原因。
英文摘要
Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these observations with a mechanistic account. We show that the model's internal computation decomposes into a four-stage sequential pipeline, Schema Abstraction, Operation Planning, Operand Binding, and Computation, each stage producing a distinct intermediate representation in an identifiable band of layers. Using the same scaffold to diagnose distractor-induced failure, we localize the corruption to a single stage, Operation Planning, implemented by a set of attention heads whose causal role we validate bidirectionally. In short, we provide a mechanistic interpretation of math word problem reasoning in LLMs, and their failure when distracted.