arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18080cs.AI

可解码性并非因果性:通过SAE分解区分探针读出与行为驱动因素

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

Devesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过SAE分解探针特征,证明探针可解码性不等于因果性,提出结合对齐与梯度敏感性的诊断方法,在真实性探针上有效分离行为驱动特征。

中文摘要 AI 辅助

线性探针可以从语言模型激活中解码诸如真实性等安全相关概念,但探针的准确性可能仅表明可解码性,而非探针权重所对应的特征在因果上驱动了模型行为。我们证明,仅凭探针权重的几何结构无法弥合这一差距:与探针方向几何对齐的特征未必是模型实际使用的特征,因此因果相关性需要干预。我们引入了一种特征级诊断方法,将已部署的真/假探针分解为稀疏自编码器(SAE)特征,根据探针对齐度和模型行为的梯度敏感性对这些特征进行排序,并在一致性门控下消融由此产生的共享特征集、仅探针特征集和随机特征集。在Buerger等人(2024)的真实性探针(TTPD)上,应用于Long等人(2025)针对Gemma2-9B-Instruct的指令性真实/欺骗设置中,两种排序仅弱重叠(约12%,Spearman rho = 0.10),且消融将它们显著分离:在完全一致性下,探针与模型共享的特征翻转输出的程度(高达27%)远大于同等规模的仅探针特征(6%)或随机特征(1%),而仅探针特征则反而扰动探针自身的读出。这种分离在五个随机种子和留出分割中均成立,且基于激活感知的特征选择翻转行为的频率几乎是探针几何顶部特征的近三倍(17.6%对比6.1%)。因此,在此设置中,探针权重向量的几何投影本身并不能识别模型因果使用的特征;然而,将探针信息与特征激活统计相结合,能恢复显著更多的行为因果特征,并且需要一致性门控的SAE干预才能将它们与探针读出区分开。

英文摘要

Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally drive model behavior. We demonstrate that this gap cannot be closed from the geometry of probe weights alone: the features geometrically aligned with probe direction need not be the ones the model uses, so causal relevance requires intervention. We introduce a feature-level diagnostic that decomposes a deployed True/False probe into sparse-autoencoder (SAE) features, ranks those features by both probe alignment and by gradient sensitivity of the model's behavior, and ablates the resulting shared, probe-only, and random feature sets under a coherence gate. On the truth probe of Buerger et al. (2024) (TTPD), applied in the instructed truth/deception setting of Long et al. (2025) for Gemma2-9B-Instruct, the two rankings overlap only weakly (about 12%, Spearman rho = 0.10), and ablation dissociates them sharply: features the probe shares with the model flip the output far more (up to 27%) than equally sized probe-only (6%) or random (1%) features at full coherence, while probe-only features instead perturb the probe's own readout. The dissociation holds across five seeds and a held-out split, and an activation-aware selection of features flips behavior nearly three times as often as the probe's geometric top features (17.6% vs. 6.1%). In this setting, therefore, the geometric projection of a probe's weight vector alone does not identify the features the model causally uses; however, combining probe information with feature activation statistics recovers substantially more behaviorally causal features, and coherence-gated SAE intervention is needed to separate them from probe readouts.

发表机构

  • WW-P High School South(西温莎-普兰斯堡高中南校)
  • Phillips Academy Andover(菲利普斯学院安多佛)
  • Dublin High School(都柏林高中)
  • Carlmont High School(卡尔蒙特高中)
  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

↑