AI 中文总结
该研究以GPT-2 small和间接宾语识别任务为对象,发现电路级可解释性证据在合理分析变异下稳定性差,多数规范对的推导声明会翻转,未通过可归档性标准。
AI 中文摘要
欧盟AI法案要求高风险系统的提供者提交技术文档,说明系统如何做出决策。机械可解释性是此类证据的明显来源,而电路发现是其最成熟的工具。我们探究此类证据是否能在其被依赖的条件下留存:两名合格分析人员、同一系统、同一工具、不同合理设置。我们预先注册了由七个分析维度组成的交叉网格,每个层级均来自已发表的实现,并通过确定性声明映射将每个发现的电路映射到结构化的附件IV声明中。在GPT-2 small和间接宾语识别任务的15840个预先注册的规范中,有7561个产生了声明,其中73.2%的规范对的推导声明发生翻转(95%置信区间为0.725至0.738),且模态声明占41.1%的空间。该证据在合格评定机构可能接受的每一个容差下均未通过可归档性标准。将最具影响力的单一选择(评估指标)标准化后,翻转率仍为59.4%。完全从声明中移除电路大小并保持固定后,翻转率为27.1%(95%置信区间为0.255至0.286),仍高于预先注册的阈值。这些声明背后的电路在结构上近乎不相交,中位数Jaccard重叠度为4%,且在Cohen's kappa为0.015时功能上不相关,因此这种不稳定性并非同一机制的不同表述。我们将可归档性标准作为独立协议给出,并报告七个已记录的发现目标之一在库自身的规范任务上根本无法执行。本研究涵盖一个模型和一个任务,结论是否适用于更大规模尚未验证。
英文摘要
The EU AI Act requires providers of high-risk systems to file technical documentation describing how the system reaches its decisions. Mechanistic interpretability is the obvious source of such evidence, and circuit discovery is its most developed instrument. We ask whether that evidence survives the condition under which it would be relied upon: two competent analysts, the same system, the same tool, different defensible settings. We pre-registered a crossed grid of seven analytic axes, every level taken from a published implementation, and mapped each discovered circuit through a deterministic claim map to a structured Annex IV statement. Across 15,840 pre-registered specifications on GPT-2 small and the indirect object identification task, of which 7,561 produced a claim, the derived statement flips across 73.2% of specification pairs (95% CI 0.725 to 0.738) and the modal claim commands 41.1% of the space. The evidence fails a filability criterion at every tolerance a conformity assessment body would plausibly accept. Standardising the single most influential choice, the evaluation metric, leaves the flip rate at 59.4%. Removing circuit size from the claim entirely and holding it fixed leaves 27.1% (95% CI 0.255 to 0.286), still above the pre-registered threshold. The circuits underlying these claims are structurally near-disjoint, median pairwise Jaccard overlap 4%, and functionally uncorrelated at Cohen's kappa 0.015, so the instability is not one mechanism described in different words. We give the filability criterion as a standalone protocol, and we report that one of the seven documented discovery objectives does not execute at all on the library's own canonical task. The study covers one model and one task, and whether the conclusion holds at scale is untested.
Comments12 pages, 1 figure, 7 tables. Pre-registered analysis plan; code and results available