arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于混合专家模型中一致专家选择的多层次上下文建模

Multi-level context Modeling for consistent expert selection in Mixture-of-Experts

Shuhan Huang, Naifan Zhang, Yuanbo Tang, Yang Li, Wai Kin Victor Chan

arXiv 2607.16427首次发表:更新:

发表机构

Tsinghua Shenzhen International Graduate School, Tsinghua University; School of AI, The Chinese University of Hong Kong (Shenzhen)(清华大学深圳国际研究生院; 香港中文大学(深圳)人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究混合专家模型中专家选择问题,提出多层次上下文融合MoE框架,通过整合跨层语义聚合和局部令牌级交互信号构建上下文感知表示,提升路由一致性和下游性能。

AI 中文摘要

混合专家模型(MoE)通过将令牌路由到一小部分专家来实现Transformer模型的高效扩展。然而,现有路由器通常基于浅层或孤立的令牌表示来进行专家选择,这往往会在各层产生不稳定且语义不一致的路由决策。在这项工作中,我们从表示角度重新审视专家选择,并将上下文不完整性识别为限制有效专家专业化的关键瓶颈。为解决此问题,我们提出了多层次上下文融合MoE(MCF-MOE)框架,该框架通过整合来自跨层语义聚合和局部令牌级交互的互补信号来构建上下文感知表示,从而实现更具信息性和一致性的专家选择。在语言建模和理解基准测试上的实验表明,MCF-MOE相对于强大的MoE基线持续提高了路由一致性和下游性能,突出了上下文完整性在专家路由中的重要性。代码可在该https URL获取。

英文摘要

Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition expert selection on shallow or isolated token representations, which often produce unstable and semantically inconsistent routing decisions across layers. In this work, we revisit expert selection from a representation perspective and identify context incompleteness as a key bottleneck limiting effective expert specialization. To address this issue, we propose Multi-level Context Fusion MOE (MCF-MOE), a framework that constructs context-aware representations by integrating complementary signals from cross-layer semantic aggregation and local token-level interactions, enabling more informative and consistent expert selection. Experiments on language modeling and understanding benchmarks demonstrate that MCF-MOE consistently improves routing consistency and downstream performance over strong MoE baselines, highlighting the importance of contextual completeness in expert routing. The code is available at https://github.com/shuhanhuang/MCF-MOE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑