AI 中文总结
本研究将稀疏MoE的专家分派与聚合解耦,提出FDAA方法,在冻结主干等模块的情况下优化聚合头,在OLMoE、DeepSeek-V2-Lite等模型上提升了语言建模性能,验证了专家选择与承诺的可区分性。
AI 中文摘要
稀疏混合专家模型(Sparse Mixture-of-Experts, MoE)的路由模块通常使用相同的分数来选择专家并对其已计算的输出进行加权。本研究探讨分派(dispatch)与聚合(aggregation)这两个角色是否应当耦合。在预训练的OLMoE-1B-7B模型上,我们保持所选Top-8专家ID、专家计算量及总选择的路由质量不变,仅改变组内聚合方式。结构化Oracle在三个随机种子下将全视野交叉熵(cross-entropy)提升0.0160±0.0039;路由模块得分最高的专家仅在17.2%的情况下是反事实最优顶点,路由效用的斯皮尔曼相关系数(Spearman)为0.030。因此,我们训练了固定分派自适应聚合(Fixed-Dispatch Adaptive Aggregation, FDAA)——一个拥有30.1万参数的后计算头,在冻结主干、路由模块及专家的同时,直接以语言建模目标进行优化。在OLMoE上,FDAA在新鲜WikiText-103测试集上的交叉熵差值(Delta CE)为-0.1523±0.0031(三个种子),混合域训练在冻结验证评估下对WikiText-103、C4及保留的Penn Treebank带来稳定增益。我们还在使用Top-6路由专家加共享专家的DeepSeek-V2-Lite上复现了固定分派审计:WikiText和C4上最优顶点的空间仍显著存在,路由Top1仅在12.5%和16.7%的审计样本中识别出所选最优专家。在单种子混合域复现中,FDAA提升了锁定的WikiText和PTB,而C4结果统计上无显著差异。这些结果支持跨架构区分专家选择与专家承诺的必要性。
英文摘要
Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study whether these two roles, dispatch and aggregation, should be coupled. On pretrained OLMoE-1B-7B, we keep selected Top-8 expert IDs, expert computation, and total selected router mass fixed and change only within-set aggregation. A structured oracle improves full-horizon cross-entropy by 0.0160 +/- 0.0039 across three seeds; the router's top-scored expert is the counterfactual-best vertex only 17.2% of the time, with router-utility Spearman 0.030. We therefore train Fixed-Dispatch Adaptive Aggregation (FDAA), a 301K-parameter post-compute head optimized directly with the language-modeling objective while freezing the backbone, router, and experts. On OLMoE, FDAA improves fresh WikiText-103 test by Delta CE = -0.1523 +/- 0.0031 across three seeds, and mixed-domain training gives robust gains on WikiText-103, C4, and held-out Penn Treebank under frozen confirmatory evaluation. We also replicate the fixed-dispatch audit on DeepSeek-V2-Lite, which uses Top-6 routed experts plus shared experts. Best-vertex headroom remains significant on WikiText and C4, while router Top1 identifies the best selected expert in only 12.5% and 16.7% of audited examples. In a one-seed mixed-domain replication, FDAA improves locked WikiText and PTB, while C4 is statistically neutral. These results support a cross-architecture distinction between expert selection and expert commitment.
Comments8 pages, 1 figure, 5 tables