SPICE:基于聚类的简单多义特征解释
SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation
- Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院(KAIST))
- INEEJI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对神经网络多义性解释的架构依赖和人工预设簇数问题,提出通用框架SPICE,自动确定簇数,首次系统比较CNN与Transformer中的多义性,并揭示其形成机制。
AI中文摘要:
神经网络可解释性近期的一个关键挑战是多义性(polysemanticity),即单个神经元被多个往往不相关的概念激活,阻碍了清晰的功能理解。尽管先前的工作已探索了这一现象,但现有方法仍局限于特定架构,并依赖人工启发式规则,如固定的概念簇数量($K$),限制了其通用性和可扩展性——尤其是对于基于Transformer的现代模型。为解决这些局限,我们提出了SPICE(\textbf{S}imple \textbf{P}olysemantic Feature \textbf{I}nterpretation via \textbf{C}lustering-based \textbf{E}xplanation,基于聚类的简单多义特征解释),一个用于分析深度视觉架构中多义性的通用框架。SPICE避免了依赖架构的传播规则,首次实现了对CNN和Transformer中多义性的系统比较,并自动确定每个神经元的概念簇数量,消除了对预设$K$的依赖,支持大规模模型的可扩展分析。利用SPICE,我们对多义性如何产生、如何随深度和架构变化以及如何通过不同的计算路径形成进行了全面研究。
英文摘要:
One of the pivotal recent challenges in neural network interpretability is polysemanticity, where a single neuron is activated by multiple, often unrelated concepts, hindering clear functional understanding. Although prior work has explored this phenomenon, existing approaches remain architecture-specific and depend on manual heuristics such as a fixed number of concept clusters ($K$), limiting their generality and scalability--especially for modern Transformer-based models. To address these limitations, we introduce SPICE (\textbf{S}imple \textbf{P}olysemantic Feature \textbf{I}nterpretation via \textbf{C}lustering-based \textbf{E}xplanation), a generalizable framework for analyzing polysemanticity in deep vision architectures. SPICE avoids architecture-dependent propagation rules, enabling the first systematic comparison of polysemanticity across both CNNs and Transformers, and automatically determines the number of concept clusters per neuron, eliminating reliance on a preset $K$ and supporting scalable analysis for large models. Using SPICE, we conduct a comprehensive investigation into how polysemanticity emerges, varies across depth and architecture, and forms through distinct computational pathways.