arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于语言、行为和机制指标检测大语言模型中的隐藏思维链

Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators

Armaan Singh, Ryan Trinh Le, Jasmine Kaur, Abdullah Sultan, Edward Lue Chee Lip, Kiran Nijjer, Adnan Ahmed, Vasu Sharma

arXiv 2608.29956首次发表:更新:

发表机构

Stanford University; York University; Facebook(斯坦福大学; 约克大学; 脸书)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出HCDS指标,通过语言、行为和机制信号检测大语言模型的隐藏思维链,在GSM8K上验证了Qwen3-4B变体存在类潜在推理行为。

AI 中文摘要

大语言模型在回答复杂推理问题时往往不会披露中间步骤,这引发了它们是在进行潜在推理还是在完成模式的疑问。我们提出了隐藏思维链检测分数(Hidden CoT Detection Score,HCDS),这是一种比较性行为与机制信号,用于测量中性提示行为与显式思维链(CoT)或显式非思维链的对齐程度。在此,隐藏思维链可操作地表示这种中性提示下的类CoT对齐;HCDS并不直接观察或证明未公开的推理轨迹。在GSM8K数据集上,两个Qwen3-4B变体的HCDS均显著为正:Thinking变体为+1.87,p值为1.2×10^-7;Instruct变体为+1.41,p值为1.9×10^-4。该结果在不同的推理栈和量化中可复现,偏差在0.08以内(分别为+1.80和+1.45),且在8个长度调整的校准对照单元中有7个未呈现显著正相关。未调整的分数在单步算术和数值事实查找任务中产生较大的正分数。这些变体对非思维链指令的反应也不同:Instruct变体仅根据提示就会服从,而Thinking变体则会继续推理并需要干预。这些发现表明,推理调优模型中存在更强、更少依赖提示的类CoT行为,这与潜在推理一致但并非其证明。因此,HCDS可在不依赖模型自我报告轨迹的情况下研究潜在推理。

英文摘要

Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete patterns. We propose the Hidden CoT Detection Score (HCDS), a comparative behavioral and mechanistic signal measuring whether neutral-prompt behavior aligns more closely with explicit CoT or explicit no- CoT. Here, hidden CoT operationally denotes this neutral-prompt CoT-like alignment; HCDS does not directly observe or prove an unexposed reasoning trace. On GSM8K, HCDS is significantly positive for both Qwen3-4B variants (Thinking $+1.87$, $p = 1.2 \times 10^{-7}$; Instruct $+1.41$, $p = 1.9 \times 10^{-4}$), replicates across a different inference stack and quantization within $0.08$ ($+1.80$ and $+1.45$), and is not significantly positive in seven of eight length-adjusted calibration-control cells. The unadjusted score produces large positive scores on single-step arithmetic and numeric factual lookup. The variants also respond differently to no-CoT instructions: Instruct complies from the prompt alone, whereas Thinking continues reasoning and requires intervention. These findings show stronger, less prompt-conditional CoT-like behavior in the reasoning-tuned model, consistent with but not proof of latent reasoning. HCDS thus investigates latent reasoning without relying on models' self-reported traces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑