arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35020cs.CV

验证线性表示假说:视觉SAE的可解释性如何?

Verifying the Linear Representation Hypothesis: How Interpretable Are Vision SAEs?

Teodor Chiaburu, Franz Motzkus, Frank Haußer, Felix Bießmann

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过适配AIS到视觉任务并开展用户研究,发现现有视觉SAE可解释性指标与AIS互不相关,证明需以真实数据为锚定来验证可解释性。

中文摘要 AI 辅助

视觉稀疏自编码器(SAE)因其据称能够将模型学习到的复杂特征解耦为单语义概念的能力,已成为机械可解释性领域的热门工具。尽管其日益流行,评估其可解释性仍是一个活跃的研究课题。采用SAE的基石是线性表示假说(LRH),该假说声称多语义特征可以投影到一组(近似)正交的、稀疏的、人类可理解的表示基上。然而,当前大多数框架评估的是诸如SAE特征的稀疏性或推断字典的连贯性等代理指标,隐含地假设这些指标反映了与人类感知的一致性。在本文中,我们提供了经验证据,表明衡量SAE概念的可解释性比这些代理指标所暗示的更为困难。为此,我们将自动可解释性评分(AIS)——先前在自然语言处理中已被证明与人类判断一致——适配到视觉任务,并在专门的用户研究中验证了我们的方法。我们使用标准指标和适配后的AIS来评估SAE概念质量。我们发现,既有的SAE可解释性指标彼此之间以及与AIS之间均不相关,这表明没有任何单一的无参考指标,无论是否基于LRH,足以验证视觉SAE的可解释性。我们认为这些发现支持了近期关于对解释方法进行更可验证、以真实数据为锚定的设计与评估的呼吁。

英文摘要

Vision Sparse Autoencoders (SAEs) have become a popular tool in Mechanistic Interpretability due to their presumed ability to disentangle complex features learned by a model into monosemantic concepts. Despite their growing popularity, evaluating their interpretability remains an active topic of research. The bedrock motivating the adoption of SAEs is the Linear Representation Hypothesis (LRH), which claims that polysemantic features can be projected onto a (near) orthogonal basis of sparse, human-understandable representations. Yet, most current frameworks evaluate proxies such as the sparsity of SAE features or the coherence of the inferred dictionary, implicitly assuming that these reflect alignment with human perception. In this paper, we provide empirical evidence that measuring the interpretability of SAE concepts is more difficult than these proxies suggest. To this end, we adapt the Autointerpretability Score (AIS) - previously shown to align with human judgments in Natural Language Processing - to vision tasks and validate our approach in a dedicated user study. We evaluate SAE concept quality using both standard metrics and our adapted AIS. We find that established interpretability metrics for SAEs correlate neither with one another nor with AIS, indicating that no single reference-free metric, whether grounded in the LRH or not, is sufficient for verifying the interpretability of vision SAEs. We argue these findings support recent calls for more verifiable, ground-truth-anchored design and evaluation of explanation methods.

发表机构

  • Fraunhofer SIT(弗劳恩霍夫安全信息技术研究所)
  • ATHENE National Research Center for Applied Cybersecurity(ATHENE国家应用网络安全研究中心)
  • AUMOVIO
  • University of Bamberg(班贝格大学)
  • Berlin University of Applied Sciences(柏林应用科学大学)
  • Einstein Center Digital Future(爱因斯坦数字未来中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑