arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAEs的独立性先验割裂了视觉概念

The Independence Prior of SAEs Fragments Visual Concepts

Tommaso Mencattini, Giorgos Nikolaou, Donato Crisostomi, Thomas Fel, Francesco Montagna, Emanuele Rodolà, Francesco Locatello

arXiv 2610.04112首次发表:更新:

发表机构

ISTA; EPFL; Goodfire; Sapienza University of Rome; Paradigma(奥地利科学技术学院; 洛桑联邦理工学院; Goodfire; 罗马第一大学; Paradigma)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对SAEs忽略视觉概念空间依赖性的问题,提出基于马尔可夫场线性表示假说的Spatial-SAE,在概念恢复与可解释性上显著优于标准SAEs。

AI 中文摘要

稀疏自编码器(SAEs)将模型激活分解为可解释字典原子的稀疏组合。尽管SAEs基于线性表示假说(LRH),但其目标函数隐含了一个额外的先验:不同patch之间的概念被视为相互独立,这一假设显然被自然图像及其所引发的激活所违背。因此,我们通过马尔可夫场线性表示假说(MFLRH)将LRH专门化到视觉领域,该假说在LRH假设中补充了缺失的空间依赖性。我们由此提出Spatial-SAE,作为MFLRH下的摊销最大后验(MAP)估计器。Spatial-SAE在概念恢复和可解释性方面持续优于标准SAEs。在四种变体上,它在合成概念恢复中实现了96%的平均胜率,并在DINOv2激活上提升了可解释性,其重建代价集中在高空间频率上。

英文摘要

Sparse Autoencoders (SAEs) decompose model activations into sparse combinations of interpretable dictionary atoms. Although SAEs are grounded in the Linear Representation Hypothesis (LRH), their objective smuggles in an additional prior: concepts across patches are treated as independent, an assumption clearly violated by natural images and by the activations they induce. We therefore specialize LRH to vision through the Markov-Field Linear Representation Hypothesis (MFLRH), which adds the missing spatial dependencies to the LRH assumptions. We thus propose Spatial-SAE as an amortized MAP estimator under the MFLRH. Spatial-SAE consistently outperforms standard SAEs in concept recovery and interpretability. Across four variants, it achieves a 96% average win rate on synthetic concept recovery and improves interpretability on DINOv2 activations, at a reconstruction cost concentrated in high spatial frequencies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑