AI 中文总结
针对语言模型自解释的机制性因果主张,提出变量特定证据标准,通过配对零假设和干扰控制检验,发现现有统计量无法确立因果敏感性,需进一步验证。
AI 中文摘要
当语言模型解释它已经给出的答案时,它是否重用了产生该答案的计算过程,还是仅从答案本身重构了一个故事?归因、可迁移性和可恢复性都与因果使用相容,但并未确立因果使用。我们提出一种证据标准:将每个正向统计量与一个变量特定的零假设配对,该零假设在尽可能匹配相关干扰维度的同时移除被测试变量的身份,并审计不匹配的维度。我们将此标准应用于一个已知原因。一个提示错误选项的线索将选择该选项的比率在三个模型中提高了64到68个百分点。在测试的四个模型中的三个,解释提及该线索的项目占比为1.8%或更低。三个估计器类别产生了有利的统计量,但在三模型分析中,没有一个在其自身控制下确立对线索对比的因果敏感性。在最强的案例中,恢复的线索方向达到$R^2$为0.95,并在所有三个种子中超过几何匹配的随机方向,而使用线索标签打乱的相同管道拟合的方向在可比较的实际编辑幅度下重现了其效果的61%到76%。第四个模型通过了一个互换端点,但不平等的编辑幅度和同时改变线索身份和线索-答案一致性的对比限制了其解释。这些实验使因果访问问题悬而未决。它们确立了一个证据要求:有利的机制统计量必须经受变量身份和干扰结构控制的检验。可重复使用的控制将通用效应与身份特定的迁移效应分开,使用打乱标签拟合零方向,并审计实际干预幅度。
英文摘要
When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each positive statistic with a variable specific null that removes the tested variable's identity while matching relevant nuisance dimensions as far as possible, and audit unmatched dimensions. We apply this standard to a known cause. A cue naming a wrong option raises the rate of choosing that option by 64 to 68 percentage points across three models. Explanations mention the cue in 1.8 percent of items or fewer in three of four models tested. Three estimator classes yield favorable statistics, but none establishes causal sensitivity to the cue contrast under its own control in the three-model analysis. In the strongest case, a recovered cue direction reaches $R^2$ of 0.95 and exceeds a geometry matched random direction in all three seeds, while a direction fitted by the same pipeline with cue labels scrambled reproduces 61 to 76 percent of its effect at comparable realized edit magnitude. A fourth model passes one interchange endpoint, but unequal edit magnitudes and a contrast that changes both cue identity and cue-answer agreement limit its interpretation. These experiments leave causal access unresolved. They establish an evidentiary requirement: favorable mechanistic statistics must survive controls for variable identity and nuisance structure. Reusable controls separate generic from identity specific transport effects, fit null directions with scrambled labels, and audit realized intervention magnitudes.