沉默的权重:潜在国际象棋推理中权重优于暂存器的因果案例
The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning
浏览论文内容
中文总结 AI 辅助
研究国际象棋模型潜在推理,通过分阶段潜在推理课程训练并强化学习,发现强化学习增加对干扰的鲁棒性而非依赖思维内容,反驳潜在思维是推理时暂存器的假设,表明其主要作用是塑造模型参数,还展示了强化学习在国际象棋领域的有效提升。
中文摘要 AI 辅助
潜在推理让语言模型在连续向量空间而非文字中进行中间计算,通常被视为模型推理时主动参考的内部暂存器。但这一假设在强化学习中是否成立尚未直接测试:现有对潜在推理的因果分析局限于数学和逻辑任务,且比较的是单个检查点内模型对其思维的依赖,而非强化学习阶段前后的情况。我们通过分阶段潜在推理课程训练国际象棋模型,随后进行强化学习,发现合法性从强化学习前的48%单调攀升至61%,同时将将杀虚构完全消除。为确定这一提升的来源,我们在强化学习前后对同一模型进行了六条件因果干预套件实验:用匹配噪声替换或添加到潜在思维向量中,性能不变;消除它们只会导致轻微下降;只有精确为零的向量会导致崩溃。这种鲁棒性差距本身就是发现:在精确归零损坏下,合法性在强化学习前降至1%,强化学习后降至9%,这一差距在对整个测试集进行校正后仍然存在;较温和的条件下趋势类似,但未独立达到显著水平。强化学习似乎增加了对干扰的鲁棒性,而非对思维内容的依赖。这些结果反驳了该领域默认的假设,即潜在思维在推理时作为主动参考的暂存器,相反表明潜在推理在此的主要作用是在训练期间塑造模型参数期间。我们还展示了在国际象棋中强化学习的有效提升,在数学和逻辑设置之外的领域,多组研究报告相同的潜在推理加强化学习方法未能提高超过监督微调的准确率。
英文摘要
Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad the model consults during inference. Whether that assumption survives reinforcement learning has not been tested directly: existing causal analyses of latent reasoning are confined to math and logic tasks, comparing reliance on thoughts within one checkpoint, never before and after RL. We train a chess-playing model through a staged latent-reasoning curriculum followed by reinforcement learning, and find legality climbs monotonically to 61% (from a 48% pre-RL baseline) while checkmate confabulation is eliminated entirely. To locate this gain, we run a six-condition causal intervention suite on the same model before and after RL: substituting or noising the thought vectors leaves performance unchanged, ablating them costs only mild degradation, and only exact-zero vectors cause collapse. This robustness gap is itself the finding: under exact-zero corruption, legality collapses to 1% pre-RL versus 9% post-RL, a gap that survives correction across the full battery. A 10x-larger replication of the post-RL checkpoint's own battery confirms this: removing the thoughts, with or without restoring sequence length, also reaches significance; substitution and noise remain indistinguishable from baseline. RL appears to add robustness to disruption, not reliance on thought content. These results push back against the field's default assumption that latent thoughts function as an actively consulted inference-time scratchpad, and instead indicate latent reasoning's principal effect here is shaping the model's parameters during training. We also demonstrate a working RL gain in chess, where multiple groups report the same latent-reasoning-plus-RL recipe failing to improve accuracy over SFT.