arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27510cs.CLcs.LG

线性探针如何产生?基于概念定向归因的电路追踪框架

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出概念定向归因(CTA)框架,通过训练针对线性探针的归因图,揭示了探针性能与可解释电路结构的关联,为探针提供机制解释并支持安全关键概念表征的审查。

中文摘要 AI 辅助

Transcoder归因图通常被训练用于解释模型为何对特定下一个token分配高概率。我们提出概念定向归因(Concept-Targeted Attribution, CTA),该方法针对线性探针方向训练归因图,从而产生特定于探针的电路,用于解释提示中内部概念表征的产生,而不依赖于其是否在生成的token中表达。我们利用跨层Transcoder(Cross-Layer Transcoders)表明,这些探针定向图包含预测性结构:图级特征可预测四个广泛研究的概念类别上的探针准确率(ρ=0.91,R²=0.84),而局部特征可识别驱动每提示分类的稀疏组件。这将探针性能与可解释的电路结构关联起来,使我们不仅能探究探针是否有效,还能探究使其有效的内部计算。因果消融进一步表明,探针定向图和logit定向图捕捉功能上不同的机制:移除探针相关特征会降低内部概念分数,同时在很大程度上保留生成的token;而移除logit相关特征会在92%至100%的案例中改变生成的token,对探针分数几乎无影响。CTA提供了一个从行为探针准确率转向探针性能的机制解释的框架,支持对内部概念表征(包括安全关键型表征)进行更详细的审查,代码可访问this https URL。

英文摘要

Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which instead trains attribution graphs with respect to a linear probe direction. CTA therefore yields probe-specific circuits that explain why an internal concept representation arises in a prompt, independently of whether it is expressed in the generated token. Using Cross-Layer Transcoders, we show that these probe-targeted graphs contain predictive structure: graph-level features predict probe accuracy across four widely studied concept categories ($ρ= 0.91$, $R^2 = 0.84$), while local features identify the sparse components driving per-prompt classification. This connects probe performance to interpretable circuit structure, allowing us to ask not only whether a probe works, but which internal computations make it work. Causal ablations further show that probe-targeted and logit-targeted graphs capture functionally distinct mechanisms. Removing probe-relevant features reduces internal concept scores while largely preserving generated tokens, whereas removing logit-relevant features changes the generated token in 92% to 100% of cases with near-zero effect on probe scores. CTA provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones. Our code is available at https://github.com/vedantpalit/concept-targeted-attribution

发表机构

  • University of Toronto(多伦多大学)
  • Vector Institute(矢量研究院)
  • Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
  • ELLIS Institute Tübingen(图宾根ELLIS研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑