arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28859cs.LGcs.AIcs.CL

停止向量:内化因果引导干预以实现高效推理

The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

Dylan Jayabahu, Tinuade Adeleke

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出通过将因果可解释性发现内化到权重中得到的停止向量,可在保持准确率的前提下减少推理模型约四分之一的思考时长,解决其不终止问题,核心贡献为该停止向量的获取方法。

中文摘要 AI 辅助

推理模型在知晓答案时不会停止。在DeepSeek-R1-Distill-Qwen-7B模型上,思维链的运行时长约为模型自身答案概率稳定所需时长的两倍,且这部分多余时长的可移除程度因问题而异,因此全局长度惩罚无法消除该多余时长。我们通过将因果可解释性发现内化到权重中来消除该多余时长。该机制是一个停止向量:此模型第18层的均值差方向,其引导强度控制模型的思考时长,而复制的值轴无作用。将该干预安装到权重中比看起来更难:最大化该方向上的标量投影会破坏冻结下游读取器所依赖的轴外维度,导致生成长度而非缩短;有效的方法是在将这些维度固定为自然值的同时重构整个引导激活。该停止向量由24个问题拟合,无需强化学习,在保持准确率的前提下,在5个未见过的基准上可移除约四分之一的思考时长,且该缩短与每个问题自身的可移除松弛度的相关性为0.70。它还能解决随难度增加而恶化的不终止问题,而解码时的置信度钩子会加剧该问题。我们不声称在原始权衡上优于调优良好的长度惩罚或解码时的提前退出;本研究的贡献在于停止向量的获取方式。

英文摘要

Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts the off-axis dimensions a frozen downstream reader depends on, and generation gets longer instead of shorter; what works is reconstructing the whole steered activation with those dimensions pinned to their natural values. Fit from 24 problems and no reinforcement learning, the halt removes about a quarter of the thinking at held accuracy across five unseen benchmarks, and the cut tracks each problem's own removable slack at 0.70. It also closes a non-termination pathology that grows with difficulty and that a decoding-time confidence hook makes worse. We do not claim to beat a well-tuned length penalty or decoding-time early exit on the raw trade-off; the contribution is how the halt is obtained.

发表机构

  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑