arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27038cs.AI

陈述的推理步骤是否具有因果承重性?

Are Stated Reasoning Steps Causally Load-Bearing?

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Abhiram Bhupatiraju, Rayan Nyaupane

中文总结 AI 辅助

本研究通过激活级因果干预测量思维链推理的因果承重性,发现行为测试高估忠实性,且模型能力影响因果忠实性。

中文摘要 AI 辅助

思维链(CoT)监控假设模型书写的推理反映了直接产生其答案的计算过程。以往的忠实性度量主要是行为层面的,因为它们只是编辑推理文本并观察由此产生的答案。然而,我们的方法旨在激活层面因果地度量忠实性,特别是针对自生成的推理。与以往测量退化程度的因果审计不同,我们的干预带有已知的预测目标。这样,每次补丁都应通过构造将答案切换到一个特定的反事实实体。具体而言,我们使用合成的多跳查找任务(2-6跳)。我们在模型陈述每个中间步骤的令牌跨度处,用来自反事实运行的相应激活来修补残差流。对于Qwen3-4B,在最敏感的网络中层,76.9% ± 2.8%的陈述步骤是因果承重的(CLB)(随机位置空值:11.3%;修补底层提示事实:83%,因此陈述步骤携带了约96%的可实现效果)。此外,对相同项的标准行为测试产生88.2%,这高估了因果忠实性11.4个百分点(项目匹配;111:14不一致对,p < 1e-15),并且对于最简单的项,高估高达20个百分点。这种差距也有明显的能力维度。Qwen3-1.7B整体上因果忠实性远低(54.8%),其忠实性随着推理深度的增加而崩溃(从2跳的68%降至6跳的30%),而Qwen3-4B保持相对平稳。尽管陈述的推理可能具有因果意义,但标准行为测试往往高估其因果忠实性,尤其是在模型推理看起来最流畅的较简单示例上。

英文摘要

Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch the answer to a specific counterfactual entity derivable by construction. Specifically, we use synthetic multi-hop lookup tasks (2-6 hops). We patch the residual stream at the token span where the model states each intermediate step with the corresponding activations from a counterfactual run. For Qwen3-4B, 76.9% +/- 2.8% of stated steps are causally load-bearing (CLB) at the most responsive mid-network layer (random-position null: 11.3%; patching the underlying prompt fact: 83%, so stated steps carry approximately 96% of the achievable effect). Moreover, the standard behavioral test on the same items yields 88.2%, which overstates causal faithfulness by 11.4 percentage points (item-matched; 111:14 discordant pairs, p < 1e-15) and, for the easiest items, by up to 20 percentage points. This gap also has a clear capability dimension. Qwen3-1.7B is far less causally faithful overall (54.8%), with its faithfulness collapsing as reasoning depth increases (68% at 2 hops to 30% at 6), while Qwen3-4B remains relatively flat. Although stated reasoning can be causally meaningful, standard behavioral tests tend to overestimate its causal faithfulness, particularly on easier examples where model reasoning appears most fluent.

补充信息

↑