arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从几何恢复到因果验证:对稀疏自编码器特征的可重复审计,从叠加几何到因果惰性

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

Mohamed Abdessalem Bal

arXiv 2607.12166首次发表:更新:

AI 中文总结

研究聚焦于稀疏自编码器特征,通过重现相关实验发现现有评估混淆不同主张。提出sae - 因果审计工具,发现大量特征存在因果惰性,重新审计细化惰性分类,应用该工具于生产SAE揭示原子碰撞信号。

AI 中文摘要

稀疏自编码器(SAEs)是将叠加的神经表示分解为可解释特征的标准,评估主要依赖于相关恢复指标——真实方向与解码器原子之间的余弦相似度。我们表明这混淆了两个不同的主张:解码器几何对齐和编码器激活行为。我们重现了Elhage等人(2022年)的叠加相图,识别出高稀疏度下的收敛伪像和极端过完备情况下描述不足的扩散共享状态。我们重现了Gao等人(2024年)的TopK与L1比较,有L1收缩的直接证据。我们的核心结果是因果性的:对每个恢复的特征进行消融和引导,我们发现在退化的SAE中高达77%通过恢复标准(余弦>=0.90)的特征——以及在训练良好的SAE中9%的特征——是因果惰性的:当特征存在时,匹配的原子从不激发,包括余弦约为1.000的匹配。我们将该方法打包为sae - 因果审计,这是一种具有确定性管道的模型无关工具。重新审计细化了发现:惰性按原因分解为结构惰性(对映体对几何,存在于良好的SAE中)和竞争惰性(退化SAE的TopK病理学),并按方向分解为读和写惰性,五个对映体对完全解离——不可监测但可通过同一个原子引导,引导特异性为143 - 310且消融效应为零。我们记录了为什么按构造无法实现字节精确的可重复性,并建议将其报告为具有明确范围的一系列主张。将该工具应用于生产SAE在小规模上重现了模式(14%惰性)并揭示了原子碰撞信号:少数原子反复成为数十个不相关概念的最接近匹配,在三个批次中重复出现。

英文摘要

Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between ground-truth directions and decoder atoms. We show this conflates two distinct claims: decoder-geometry alignment and encoder-activation behavior. We reproduce the superposition phase diagram of Elhage et al. (2022), identifying a convergence artifact at high sparsity and an under-described diffuse sharing regime at extreme overcompleteness. We reproduce the TopK-versus-L1 comparison of Gao et al. (2024), with direct evidence of L1 shrinkage. Our central result is causal: subjecting every recovered feature to ablation and steering, we find up to 77% of features passing a recovery bar (cosine >= 0.90) in a degraded SAE -- and 9% in a well-trained one -- are causally inert: the matched atom never fires when the feature is present, including matches at cosine ~1.000. We package the method as sae-causal-audit, a model-agnostic instrument with a deterministic pipeline. Re-auditing refines the finding: inertness decomposes by cause into structural inertness (antipodal-pair geometry, present in good SAEs) and competitive inertness (a TopK pathology of degraded SAEs), and by direction into read- and write-inertness, which five antipodal pairs dissociate completely -- unmonitorable yet steerable through the same atom, with steering specificities of 143-310 attached to zero ablation effects. We document why byte-exact reproducibility is unavailable by construction, and propose reporting it as a stack of claims with explicit scopes. Applying the instrument to a production SAE reproduces the pattern at small scale (14% inert) and surfaces an atom-collision signal: a handful of atoms recur as the nearest match for dozens of unrelated concepts, replicated across three batches.

Comments21 pages, 8 figures, 6 tables. Code and reproduction pipeline: https://github.com/mohamed-bal/sae-causal-audit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑