发表机构
University of Maryland; SAP Labs, LLC(马里兰大学; 思爱普实验室有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究思维链模型的双峰收敛模式,通过在特定令牌位置的隐藏状态激活上训练线性探针,发现收敛命运可在生成结束前于中间表征中部分编码,为早期退出推理和自适应计算分配开辟道路。
AI 中文摘要
思维链推理模型如DeepSeek-R1-Distill-Qwen-7B呈现双峰收敛模式:生成在令牌预算内终止(收敛)或耗尽预算未得出结论(未收敛)。我们通过实证表征此现象,表明收敛生成在AIME 1983 - 2024上准确率达90.3%,未收敛的仅6.6%,总体收敛率62.0%。接着研究能否利用模型内部表征在思维链早期检测此结果。在令牌位置50 - 300的隐藏状态激活上训练线性探针,发现令牌150处第20层激活的AUC为0.608,即使在令牌50处也可靠地高于随机水平。激活探针始终优于基于令牌熵和重复统计的行为基线。扫描级排列检验得p = 0.063。这些发现表明收敛命运在生成结束前就部分编码在中间表征中。
英文摘要
Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion (non-converged). We characterize this phenomenon empirically, showing that converged generations achieve 90.3% accuracy on AIME 1983-2024 while non-converged ones achieve only 6.6%, with an overall convergence rate of 62.0%. We then ask whether this outcome is detectable early in the thinking chain using internal model representations. Training linear probes on hidden-state activations at token positions 50-300, we find that layer-20 activations at token 150 achieve AUC 0.608 (+-0.080, 5-fold CV), reliably above chance even at token 50. Activation probes consistently outperform behavioral baselines derived from token entropy and repetition statistics. A sweep-level permutation test yields p=0.063 (100,000 permutations), consistent with a modest signal that our sample size cannot confirm at conventional thresholds. These findings suggest that convergence fate is partially encoded in intermediate representations well before the generation ends, opening a path toward early-exit inference and adaptive compute allocation.