发表机构
Stanford University; York University; Facebook(斯坦福大学; 约克大学; 脸书)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出HCDS指标,通过语言、行为和机制信号检测大语言模型的隐藏思维链,在GSM8K上验证了Qwen3-4B变体存在类潜在推理行为。
AI 中文摘要
大语言模型在回答复杂推理问题时往往不会披露中间步骤,这引发了它们是在进行潜在推理还是在完成模式的疑问。我们提出了隐藏思维链检测分数(Hidden CoT Detection Score,HCDS),这是一种比较性行为与机制信号,用于测量中性提示行为与显式思维链(CoT)或显式非思维链的对齐程度。在此,隐藏思维链可操作地表示这种中性提示下的类CoT对齐;HCDS并不直接观察或证明未公开的推理轨迹。在GSM8K数据集上,两个Qwen3-4B变体的HCDS均显著为正:Thinking变体为+1.87,p值为1.2×10^-7;Instruct变体为+1.41,p值为1.9×10^-4。该结果在不同的推理栈和量化中可复现,偏差在0.08以内(分别为+1.80和+1.45),且在8个长度调整的校准对照单元中有7个未呈现显著正相关。未调整的分数在单步算术和数值事实查找任务中产生较大的正分数。这些变体对非思维链指令的反应也不同:Instruct变体仅根据提示就会服从,而Thinking变体则会继续推理并需要干预。这些发现表明,推理调优模型中存在更强、更少依赖提示的类CoT行为,这与潜在推理一致但并非其证明。因此,HCDS可在不依赖模型自我报告轨迹的情况下研究潜在推理。
英文摘要
Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete patterns. We propose the Hidden CoT Detection Score (HCDS), a comparative behavioral and mechanistic signal measuring whether neutral-prompt behavior aligns more closely with explicit CoT or explicit no- CoT. Here, hidden CoT operationally denotes this neutral-prompt CoT-like alignment; HCDS does not directly observe or prove an unexposed reasoning trace. On GSM8K, HCDS is significantly positive for both Qwen3-4B variants (Thinking $+1.87$, $p = 1.2 \times 10^{-7}$; Instruct $+1.41$, $p = 1.9 \times 10^{-4}$), replicates across a different inference stack and quantization within $0.08$ ($+1.80$ and $+1.45$), and is not significantly positive in seven of eight length-adjusted calibration-control cells. The unadjusted score produces large positive scores on single-step arithmetic and numeric factual lookup. The variants also respond differently to no-CoT instructions: Instruct complies from the prompt alone, whereas Thinking continues reasoning and requires intervention. These findings show stronger, less prompt-conditional CoT-like behavior in the reasoning-tuned model, consistent with but not proof of latent reasoning. HCDS thus investigates latent reasoning without relying on models' self-reported traces.