AI 中文总结
本研究将Transformer视为浅层电路,提出任务极限概念,证明串行任务中思维链不可或缺,并通过有限群文字题及MATH-500、AIME实验验证了擦除思维链会导致性能显著下降。
AI 中文摘要
推理模型在回答之前会生成很长的思维链,但这些链的内容是否真正完成了计算工作,还是很大程度上只是装饰性的,目前仍存在争议。我们通过将 Transformer 视为浅层电路来研究这一问题。通过固定数量层的一次前向传播具有恒定深度,因此任何运行模型恒定次数的过程都是一个浅层电路。我们将浅层电路在某个任务上所能达到的最佳准确率称为该任务的极限,如果某个任务的极限低于 1,则该任务是串行的。我们证明了关于串行任务的三个结果,这些结果对每个 Transformer 都成立,无论其如何训练。必要性:用任何不依赖于其内容的东西(例如填充标记或问题的重述)替换思维链,会将准确率降至极限,而在最大串行任务上则会降至随机水平。深度:没有浅层计算能够写出准确率超过极限的模型的思维链,甚至近似都做不到。局部性:答案与完成的思维链只相隔一次浅层传递,因此所有串行推理都发生在思维链中。在有限群的文字题上(其极限已知),从头训练的小型 Transformer,无论是否使用强化学习,都达到了预测的数字:经过思维链训练的模型能解决所有输入长度的问题,当思维链被擦除时准确率降至随机水平;没有思维链的模型随着输入长度增长准确率降至极限;开放权重推理模型在给定相同文字问题时,如果没有思维链则回到基线水平。在 MATH-500 和 AIME 上,擦除思维链会使开放推理模型的准确率降低 0.52 到 0.82,句子打乱是无害的,而标记打乱与擦除同样有害;对于使用正确或随机奖励的 GRPO 训练的检查点也是如此。因此,任务的极限回答了 Transformer 何时能在没有思维链的情况下成功。
英文摘要
Reasoning models generate long chains of thought before they answer, yet it is debated whether the content of these chains does real computational work or is largely decorative. We study this question by viewing a transformer as a shallow circuit. One forward pass through a fixed number of layers has constant depth, so any procedure that runs the model a constant number of times is a shallow circuit. We call the best accuracy that a shallow circuit can reach on a task the ceiling of the task, and a task is serial if its ceiling lies below one. We prove three results on serial tasks that hold for every transformer, no matter how it was trained. Necessity: replacing the chain by anything that does not depend on its content, such as filler tokens or a restatement of the question, drives the accuracy down to the ceiling, and on a maximally serial task down to chance. Depth: no shallow computation can write the chain of a model whose accuracy exceeds the ceiling, not even approximately. Locality: the answer is one shallow pass away from the finished chain, so all of the serial reasoning happens in the chain. On word problems of finite groups, whose ceilings are known, small transformers trained from scratch, with or without reinforcement learning, attain the predicted numbers: chain-trained models solve every input length and fall to chance when the chain is erased, chainless models collapse to the ceiling as the input length grows, and open-weight reasoning models given the same problem in words return to the baseline without their chain. On MATH-500 and AIME, erasing the chain costs open reasoning models 0.52 to 0.82 accuracy, a sentence shuffle is harmless, and a token shuffle is as harmful as erasing; the same holds for checkpoints trained by GRPO with a correct or a random reward. The ceiling of a task therefore answers when a transformer can succeed without its chain of thought.