arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.07639cs.LGcs.CL

使专家混合模型具备内在可解释性

Mixture of Experts Made Intrinsically Interpretable

  • University of Oxford(牛津大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingyi Yang, Constantin Venhoff, Ashkan Khakzar, Christian Schroeder de Witt, Puneet K. Dokania, Adel Bibi, Philip Torr

更新

AI总结:

本文提出MoE-X,通过将专家混合层重构为稀疏大型MLP、在专家内部强制稀疏激活并改进路由机制,使语言模型在保持性能的同时具备内在可解释性,可解释性超越稀疏自编码器方法。

AI中文摘要:

大型语言模型中的神经元常常表现出多义性,即同时编码多个不相关的概念,从而掩盖了可解释性。我们不依赖事后方法,而是提出了MoE-X,这是一种专家混合(MoE)语言模型,旨在具备内在可解释性。我们的方法源于以下观察:在语言模型中,具有稀疏激活的更宽网络更有可能捕获可解释的因素。然而,直接训练这种大型稀疏网络在计算上是难以承受的。MoE架构通过仅对任意给定输入激活一部分专家,提供了一种可扩展的替代方案,并且天然地与可解释性目标相一致。在MoE-X中,我们通过将MoE层重写为等效的稀疏大型MLP来建立这种联系。这种方法能够在保持稀疏性的同时高效扩展隐藏层大小。为了进一步增强可解释性,我们在每个专家内部强制实现稀疏激活,并重新设计路由机制,以优先选择激活稀疏度最高的专家。这些设计确保只有最显著的特征才会被路由并由专家处理。我们在国际象棋和自然语言任务上评估了MoE-X,结果表明其性能可与稠密模型相媲美,同时显著提高了可解释性。MoE-X实现了优于GPT-2的困惑度,其可解释性甚至超过了基于稀疏自编码器(SAE)的方法。

英文摘要:

Neurons in large language models often exhibit \emph{polysemanticity}, simultaneously encoding multiple unrelated concepts and obscuring interpretability. Instead of relying on post-hoc methods, we present \textbf{MoE-X}, a Mixture-of-Experts (MoE) language model designed to be \emph{intrinsically} interpretable. Our approach is motivated by the observation that, in language models, wider networks with sparse activations are more likely to capture interpretable factors. However, directly training such large sparse networks is computationally prohibitive. MoE architectures offer a scalable alternative by activating only a subset of experts for any given input, inherently aligning with interpretability objectives. In MoE-X, we establish this connection by rewriting the MoE layer as an equivalent sparse, large MLP. This approach enables efficient scaling of the hidden size while maintaining sparsity. To further enhance interpretability, we enforce sparse activation within each expert and redesign the routing mechanism to prioritize experts with the highest activation sparsity. These designs ensure that only the most salient features are routed and processed by the experts. We evaluate MoE-X on chess and natural language tasks, showing that it achieves performance comparable to dense models while significantly improving interpretability. MoE-X achieves a perplexity better than GPT-2, with interpretability surpassing even sparse autoencoder (SAE)-based approaches.

↑