每个 token 都在思维流中留下涟漪:引出模型内部 token 显著性以实现思维链压缩
Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression
- University of Virginia(弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对思维链推理成本过高的问题,提出模型内部显著性方法 MIST,沿必要性与充分性轴定义 token 重要性并剪枝,在四个基准和四个模型上优于基线。
AI中文摘要:
思维链(CoT)推理可提升多步骤问题解决能力,但较长的推理轨迹会增加推理成本。token 级 CoT 压缩通过将完整推理链剪枝为更短的轨迹以适配模型,使得 token 选择成为核心挑战。现有方法通常仅依赖与模型内部答案计算间接相关的外部评分器或启发式信号。我们转而采用模型内部视角:当模型生成答案时,每个推理 token 都会在残差流(即模型的“思维流”)中留下一个涟漪,该涟漪的幅度反映了 token 对答案计算的贡献。基于这一观点,我们提出 \textsc{MIST}(用于 token 级 CoT 压缩的模型内部显著性),其沿两个互补轴定义 token 重要性:必要性,即移除某 token 的内部贡献时答案似然的下降幅度;充分性,即仅提供该贡献时答案似然的提升幅度。将两者结合可得到用于剪枝的统一重要性分数。在四个推理基准和四个模型上,\textsc{MIST} 始终优于基线方法,表明模型内部显著性是推理 token 重要性的有效替代指标。
英文摘要:
Chain-of-thought (CoT) reasoning improves multi-step problem solving, but long reasoning traces inflate inference cost. Token-level CoT compression reduces this cost by pruning full reasoning chains into shorter traces for model adaptation, making token selection the central challenge. Existing methods often rely on external scorers or heuristic signals only indirectly tied to the model's internal answer computation. We instead adopt a model-internal perspective: as the model forms an answer, each reasoning token induces a ripple in the residual stream whose effect on the answer reflects the token's contribution to the underlying computation. Building on this view, we propose \textsc{MIST} (Model-Internal Saliency for Token-level CoT compression), which defines token importance along two complementary axes: \emph{necessity}, the drop in answer likelihood when a token's internal contribution is removed, and \emph{sufficiency}, the gain in answer likelihood when that contribution alone is provided. Combining the two yields a unified importance score for pruning. Across four reasoning benchmarks and four models, \textsc{MIST} consistently outperforms baseline methods, suggesting that model-internal saliency provides an effective proxy for reasoning-token importance.