发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对EEG基础模型中SAE特征解释的误判风险,提出验证阶梯方法,通过多级检验确认特征语义,并证明仅凭扰动敏感性不足以确立可解释性。
AI 中文摘要
稀疏自编码器(SAEs)将密集的模型激活分解为离散的潜在变量,使得单个特征易于解释——也易于误解。在EEG基础模型中,这产生了一个诱人的推断:如果移除alpha频带活动强烈改变潜在变量的激活,人们可能会得出结论,该潜在变量代表alpha活动。在跨越三个骨干网络、三个EEG数据集和三个网络深度的27种设置中,这种解释最初似乎令人信服:alpha移除改变潜在激活的程度是等宽假切迹(sham notch)的7.3倍(95%置信区间[6.2, 8.7],基于设置的自举法)。然而,alpha滤波器也删除了比假切迹多得多的信号。在按移除的频谱能量归一化后,比率降至0.28(95%置信区间[0.22, 0.36]),并且在27种设置中没有一个超过1。针对alpha移除响应而选择的潜在变量,在干净的EEG上,与相对alpha功率略微负相关(平均r = -0.073),这为简单的alpha检测器解读提供了不支持。受此失败案例的启发,我们提出了一个用于SAE潜在变量语义解释的验证阶梯:它依次询问潜在变量是否响应,该响应是否在控制干预移除的信号量后仍然存在,是否具有特异性而非广泛脆弱性,以及所提出的属性是否在未扰动数据上可见——同时单独测试更强的声明,即该潜在变量对任务分类器是否重要。仅扰动敏感性并不能确定SAE潜在变量代表什么。
英文摘要
Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly changes a latent's activation, one might conclude that the latent represents alpha activity. Across 27 settings spanning three backbones, three EEG datasets, and three network depths, this interpretation initially appears compelling: alpha removal changes latent firing 7.3 times more than an equal-width sham notch (95% CI [6.2, 8.7], bootstrapped over settings). However, the alpha filter also deletes far more signal than the sham. After normalizing by removed spectral energy, the ratio falls to 0.28 (95% CI [0.22, 0.36]) and exceeds one in none of the 27 settings. Latents selected for their response to alpha removal are, on clean EEG, slightly anti-correlated with relative alpha power (mean r = -0.073), giving no support for a simple alpha-detector reading. Motivated by this failure case, we propose a validation ladder for semantic interpretations of SAE latents: it asks in turn whether a latent responds, whether that response survives controlling for how much signal the intervention removes, whether it is specific rather than broadly fragile, and whether the proposed property is visible on unperturbed data--while separately testing the stronger claim that the latent matters to a task classifier. Perturbation sensitivity alone does not establish what an SAE latent represents.