arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在信任可解释性工具的裁决之前校准它们

Calibrating Interpretability Instruments Before Trusting Their Verdicts

Orion Reblitz-Richardson

arXiv 2609.14754首次发表:更新:

发表机构

Distiller Labs(Distiller 实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文识别并系统化了LLM可解释性测量中的六种失败模式,提出校准、验证、功效计算和深度参照四步纪律,以确保因果论断的可靠性。

AI 中文摘要

关于大型语言模型(LLM)内部机制的因果性论断依赖于测量。这些测量可能包括投影、余弦相似度、消融差异或交换补丁等。这些测量会以特定的、可诊断的方式失败,返回一个看似合理的数字而不是错误,因此一个损坏的工具很容易被误认为是一个发现。协方差匹配的零假设可能饱和,直到每个方向看起来都典型;在重归一化架构上,每头归因可能超出真实残差写入的三倍;交换补丁可能因为其结果被固定在某个上限而变得符号混乱;或者读取裁决可能是测量超过模型已经做出决定的层之后的产物。本文记录了来自一个关于拒绝和道德表征的因果可解释性项目中的六种此类失败,该项目涵盖多篇论文和一个四模型开放权重面板;每种模式在四个模型中的一两个上得到证实。对于每种模式,我们给出了能够捕捉它的迹象和一个针对可检测触发条件(重归一化、大规模激活、低维决策通道)的协议,以便我们和读者能够检查给定的设置是否暴露于这些风险。这一纪律归结为四个步骤:对照阳性对照阶梯进行校准,用正交单元进行验证,在花费计算之前计算统计功效,并在参考模型承诺的深度上陈述每个读取裁决。证据是单一项目内三个家族中的四种架构;跨项目的外部复制是未来的工作。

英文摘要

Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an interchange patch among others. These measurements fail in specific, diagnosable ways that return a plausible number instead of an error, so a broken instrument can easily read as a finding. A covariance-matched null can saturate until every direction looks typical, a per-head attribution can overshoot the true residual write threefold on reordered-normalization architectures, an interchange patch can go sign-chaotic because its outcome is pinned at a ceiling, or a read-from verdict can be an artifact of measuring past the layer where the model already decided. This note documents six such failures from a causal interpretability program on refusal and moral representation, spanning several papers and a four-model open-weight panel; each mode is established on one or two of the four. For each we give the tell that catches it and a protocol keyed to a detectable trigger (reordered normalization, massive activations, a low-dimensional decision channel), so we and readers can check whether a given setup is exposed. The discipline reduces to four moves: calibrate against a positive-control ladder, certify with an orthogonal cell, compute power before spending compute, and state every read-from verdict at a depth referenced to the model's commitment. The evidence is four architectures across three families within a single program; external replication across programs is future work.

Comments16 pages, 3 figures, 1 table. Companion to "Refusal Reads Only a Slice of What the Model Knows" (submitted concurrently). Per-unit arrays: https://doi.org/10.5281/zenodo.22731361. Code: https://github.com/deepsteer/deepsteer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑