AI 中文总结
该研究针对可解释人工智能中评估稀疏自编码器单语义性的挑战,提出无标签的Tversky单语义性分数(TMS),通过激活集一致性衡量。在多种模型和设置下评估,结果显示TMS受编码器各向异性影响小,能揭示训练动态,还能体现概念删除有效性。
AI 中文摘要
在可解释人工智能中,机械可解释性利用稀疏自编码器(SAE)从神经表示中提取更具可解释性的特征。然而,评估它们的单语义性以及解释质量仍然具有挑战性。现有指标需要外部概念标签或依赖预训练的嵌入模型,这使得它们对编码器的几何结构敏感。我们引入了Tversky单语义性分数(TMS),这是一种无标签指标,将单语义性作为二值化SAE潜在的激活集一致性来操作,并且不需要外部嵌入编码器。我们在基于预训练视觉和视觉语言模型(DINOv3、CLIP、BLIP2)的特征训练的SAE上评估TMS,包括两种常见的SAE模式(TopK、BatchTopK)、多个稀疏度水平和扩展因子。我们的结果表明,TMS比基于嵌入的替代方法受编码器各向异性的影响更小,同时与已建立的单语义性指标保持一致。TMS还揭示了不同基础模型上SAE的不同训练动态。此外,在编码器各向异性下,TMS能更有力地表明基于探针的概念删除有效性,在其他方面也具有竞争力。
英文摘要
Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations. However, assessing their monosemanticity, and thus explanation quality, remains challenging. Existing metrics require external concept labels or depend on pretrained embedding models, making them sensitive to encoder's geometry. We introduce the Tversky Monosemanticity Score (TMS), a label-free metric that operationalizes monosemanticity as activation-set coherence of binarized SAE latents, and does not require external embedding encoders. We evaluate TMS on SAEs trained on features from pretrained vision and vision-language models (DINOv3, CLIP, BLIP2), two common SAE regimes (TopK, BatchTopK), multiple sparsity levels, and expansion factors. Our results show that TMS is less affected by encoder anisotropy than its embedding-based alternative, while remaining aligned with established monosemanticity indicators. TMS also reveals distinct SAE training dynamics across base models. Moreover, under encoder anisotropy, TMS provides a stronger indication of probe-based concept deletion effectiveness, while being competitive otherwise.
CommentsThis is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper