arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越几何互补性:稀疏混合专家路由中的相干重叠

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

Huiyuan Tian, Bonan Xu, Shijian Li

arXiv 2607.28308首次发表:更新:

AI 中文总结

该研究区分了MoE路由的相干性等量,提出相干重叠概念,发现所选专家子空间重叠显著但实际路由更优,且添加后续专家多能提升预测,说明几何相似性无法决定冗余或剪枝价值。

AI 中文摘要

稀疏混合专家(MoE)语言模型将每个token路由至多个专家,这表明其优势可从几何角度解释:被共同选择的专家应贡献不同的表示方向。现有证据常混淆路由相干性、候选质量及候选-上下文交互。我们利用专家子空间分离指数(ESSI)、匹配路由残差及前缀控制的2×2析因设计区分这些量;冻结路由干预及受控Top-k研究评估功能价值。三项配对对比组织了研究结果:第一,在6种MoE架构中,专家子空间重叠显著,但实际路由对token表示的解释优于匹配替代方案;第二,在OLMoE、Mixtral及DeepSeek的39个析因单元中,所选候选对残差表示的解释优于每个单元中最强的未选对手,但实际前缀全程缩小了这一优势:所有交互均为负,且每个95%置信区间均低于零;第三,这种几何缩小并不意味着功能冗余:在39项冻结路由对比中,添加后续专家在24项中改善了下一个token预测,其余15项估计无定论;受控训练研究在所有3个随机种子中也支持Top-2优于Top-1。我们将此联合模式称为相干重叠:路由从共享几何邻域中选择与token相关的专家,同时有用的多专家计算在无离散线性覆盖的情况下持续存在。区分这些量阐明了为何仅几何相似性无法决定冗余或剪枝价值。

英文摘要

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled $2\times2$ factorial; frozen-route interventions and a controlled Top-$k$ study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑