arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RouterInterp:理解混合专家模型路由中的叠加式专业化

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde

arXiv 2610.11775首次发表:更新:

发表机构

Faculty of Informatics, TU Wien; UK AI Security Institute(维也纳工业大学信息学学院; 英国人工智能安全机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对混合专家模型路由的不可解释问题,提出叠加式专业化假设,开发RouterInterp方法,在gpt-oss-20b上使路由解释检测准确率较旧方法提升约65%,增进了对基础模型该组件的理解。

AI 中文摘要

稀疏混合专家(MoE)模型通过将token路由至仅对部分token激活的模块化专家网络,实现了比密集模型更高效的扩展。关于MoE模型性能的主流假设是,每个专家都专攻单一、连贯的领域。然而,基于该假设开展的可解释性研究普遍未取得成功。我们提出并为一种被称为叠加式专业化假设(SSH)的替代解释提供了证据:专家专攻的是细粒度特征的不相交并集,而非单一宽泛领域。利用SSH,我们引入了RouterInterp,这是一种用于解释专家路由的方法,可识别对路由决策最具预测性的稀疏自编码器特征,并生成统一的自然语言解释。在gpt-oss-20b上,RouterInterp解释专家路由的检测准确率比以往基于token统计的方法高出约65%。本研究提供了一种可扩展的方法,用于生成更准确的专家路由解释,并增进了我们对基础模型这一此前不可解释组件的理解。

英文摘要

Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with ${\sim}65\%$ higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.

Comments33 pages (12 non-appendix pages), 7 figures, published as a conference paper at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑