arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

0.6 有多高?可解释性探测中的下限、上限与可提升空间

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

Pranjal Garg

arXiv 2610.08544首次发表:更新:

AI 中文总结

针对探测分数缺乏固定含义的问题,提出以输入下限和上限为参考点计算可提升空间,证明其在两种条件下消失,并在元分析 Transformer、scGPT 及四项 LLM 探测研究中验证,部分声称可被文本本身解释。

AI 中文摘要

探测是可解释性研究的主力工具。如果模型隐藏状态能够预测某个变量,就认为模型表征了该变量。但探测分数没有固定含义。R² 为 0.6 可能仅仅反映了输入本身已泄露的信息,而且相同的分数在不同数据上可能意味着不同的事情。我们建议将每个探测分数与两个参考点对照:下限,即一组声明的简单输入本身已能预测的内容;以及上限,即完整输入所能预测的内容。两者之间的差距,即可提升空间,是探测能够展示模型在简单输入之外还计算了更多内容的范围。我们证明可提升空间会以两种方式消失:目标变量不再依赖于模型必须推断的隐藏变量,或者输入不再揭示该隐藏变量。我们在为上下文元分析训练的 Transformer 上测试了这一点,这些模型必须推断研究之间的隐藏异质性以正确加权,且这两个参考点均为已知。在分布偏移下,探测分数下降,预测误差上升 12 至 15 倍,但模型恢复了相似比例的可提升空间,表明是数据丢失了信息,而非表征丢失了信息。随后我们分析了真实模型。单细胞基础模型 scGPT 仅部分编码了生物变异性。我们还重新审视了四项有影响力的 LLM 探测研究,这些研究声称模型表征了地理、奥赛罗棋盘状态、真实性及其用户的人口统计特征。与仅从输入文本计算出的下限相比,其中一些主张成立,而另一些则很大程度上可由文本本身解释。

英文摘要

Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises $12$--$15\times$, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑