发表机构
IvLabs, VNIT Nagpur(伊夫实验室,维什瓦拉亚国立理工学院那格浦尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TopoLoss作为训练时先验,在ViT上通过空间局部性损失集中因果回路,实验表明其提升因果充分性2.79倍且不改变神经元单语义性,揭示现有指标盲区。
AI 中文摘要
视觉Transformer的机制可解释性旨在将模型计算分解为人类可读的单元,但学习到的表示在每个神经元中纠缠了众多概念。特征叠加被广泛视为这一分解的核心障碍,然而大多数缓解方法(稀疏自编码器、字典学习)都是事后性的,且不改变底层网络。我们探究空间局部性训练损失(TopoLoss)能否作为一种轻量级的训练时先验,提升标准机制可解释性工具的可解释性。在ImageNet-100上以多个TopoLoss权重$\alpha$训练ViT,我们通过激活修补测量拓扑簇的因果充分性,并通过拟合到同一残差流的稀疏自编码器测量特征几何。在$\alpha=1.0$时,拓扑簇的因果充分性比同尺寸随机单元集高2.79倍,且该效应随$\alpha$单调递增。SAE的L0稀疏性降低11%,死特征比例上升19倍,但标准神经元级单语义性分数不变,表明拓扑压力在回路层面起作用,将因果质量集中到空间局部结构中而不解开单个神经元。这种分离表明当前神经元级单语义性指标对一类真实可解释性增益不敏感,并将廉价的架构先验定位为事后工具可行的训练时补充。
英文摘要
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposition is widely treated as the central obstacle to this decomposition, yet most mitigations (sparse autoencoders, dictionary learning) are post-hoc and leave the underlying network unchanged. We ask whether a spatial-locality training loss (TopoLoss) can act as a lightweight, training-time prior that improves interpretability of standard mech-interp tools. Training ViT on ImageNet-100 across multiple TopoLoss weights $α$, we measure causal sufficiency of topographic clusters via activation patching and feature geometry via sparse autoencoders fit to the same residual stream. At $α=1.0$, topographic clusters are 2.79$\times$ more causally sufficient than random unit sets of the same size, with the effect increasing monotonically in $α$. SAE L0 sparsity decreases by 11% and dead-feature fraction rises 19-fold, yet standard neuron-level monosemanticity scores are unchanged, indicating that topographic pressure acts at circuit level, concentrating causal mass into spatially local structures without disentangling individual neurons. This dissociation suggests current neuron-level monosemanticity metrics are insensitive to a class of real interpretability gains, and positions cheap architectural priors as a viable training-time complement to post-hoc tooling.
CommentsAccepted at the Mechanistic Interpretability Workshop at ICML 2026