arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合专家大语言模型中的交叉熵引导路由

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit

arXiv 2609.37751首次发表:更新:

发表机构

Bar-Ilan University; NVIDIA(巴伊兰大学; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对稀疏混合专家模型路由与词元错误不对齐的问题,提出两种基于交叉熵对齐的词元错误监督机制,在多个基准上提升准确率约2.3个百分点。

AI 中文摘要

稀疏混合专家(MoE)大语言模型通过将每个词元路由到一小部分专家来扩展模型容量。其路由器通过负载均衡项进行正则化,并通过语言模型目标学习亲和度分数。然而,这些目标并未在路由亲和度与词元级错误之间提供直接对齐。我们以两种形式引入用于稀疏路由的词元错误监督。第一种形式为每个专家预测一个错误分数。这些分数的亲和度加权聚合与下一词元交叉熵损失对齐,而各个分数在top-K选择之前对亲和度进行衰减。第二种形式直接将路由器的亲和度与模型目标对齐,无需额外的头或推理时修改。两种形式均使用Itakura-Saito散度或指数负对数似然来对齐亲和度与词元错误。在两个稀疏MoE骨干网络和四个多项选择问答基准上,我们评估了这两种监督机制。在Granite上,与参数匹配的路由基线相比,我们的方法平均准确率提高了约2.3个百分点。在更强的监督下,ARC-Challenge上的增益达到2.94个百分点。两种机制均保留了原生稀疏执行预算和聚合策略。我们的代码可在补充材料中获取。

英文摘要

Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.

Comments25 pages, 6 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑