arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从概念对齐到因果锚定:思维链忠实性的干预测试

From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

Qianli Wang, Yilong Wang, Dennis Wei, Jingyi Sun, Simon Ostermann, Isabelle Augenstein, Pepa Atanasova, Nils Feldhus

arXiv 2609.23065首次发表:更新:

发表机构

German Research Center for Artificial Intelligence (DFKI); Saarland Informatics Campus; Centre for European Research in Trusted AI (CERTAIN); Technische Universität Berlin; IBM Research; University of Copenhagen; University of Groningen(德国人工智能研究中心(DFKI); 萨尔兰信息学园区; 欧洲可信人工智能研究中心(CERTAIN); 柏林工业大学; IBM研究院; 哥本哈根大学; 格罗宁根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将思维链忠实性定义为内部概念锚定,通过共享稀疏自编码器和因果指标Δp,发现概念对齐高但因果贡献随层深变化,且因果概念未必被言语化,强调需因果测试评估忠实性。

AI 中文摘要

思维链(CoT)可以听起来合理,却可能对模型潜在的推理过程不忠实。以往大多数工作通过输入-输出行为或输入归因来探测CoT的忠实性,对内部计算的研究则相对不足。我们转而将忠实性视为内部概念锚定:大语言模型(LLM)的CoT推理是否涉及与支持LLM直接预测相同的内部概念,并且这些共享概念是否因果地驱动其答案?通过使用一个共享的稀疏自编码器(SAE)对预测通道和CoT通道进行编码,SAE是LLM使用的潜在概念的可靠近似器,从而使它们的内部概念可以直接比较。我们引入了三个概念级对齐的相关性指标和一个因果指标Δp,该指标通过消融共享概念并测量答案概率的下降来评估因果贡献。在五个LLM和四个数据集上,概念对齐普遍较高,如相关性指标所示;然而这些指标仅识别哪些概念是共享的,并未量化它们的因果贡献程度。Δp填补了这一空白:因果忠实性随模型深度显著变化,在中后层而非最终层达到峰值,并且模型规模重塑了逐层的分布特征。此外,因果重要的共享概念并不总是在CoT中被言语化。这些分离现象表明,仅凭表面或表征对应关系无法可靠评估忠实性;评估它需要对CoT背后的内部概念是否真正驱动模型预测进行因果测试。

英文摘要

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large language model's (LLM) CoT reasoning engage the same internal concepts that support the LLM's direct prediction, and do the shared concepts causally drive its answer? Encoding a prediction pass and a CoT pass with a single shared sparse autoencoder (SAE), a reliable approximator of the latent concepts LLMs use, makes their internal concepts directly comparable. We introduce three correlational metrics of concept-level alignment and a causal metric, $Δp$, which ablates the shared concepts and measures the drop in answer probability. Across five LLMs and four datasets, concept alignment is generally high, as indicated by the correlational metrics; yet these only identify which concepts are shared, not how much they causally contribute. $Δp$ fills this gap: causal faithfulness varies substantially with model depth, peaking at mid-to-late layers rather than the final ones, and model scale reshapes the layer-wise profile. Moreover, causally important shared concepts are not always verbalized in the CoT. These dissociations suggest that faithfulness cannot be reliably assessed from surface-level or representational correspondence alone; assessing it requires causal tests of whether the internal concepts underlying a CoT actually drive the model's prediction.

CommentsIn submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑