arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏概率映射如何塑造混合专家路由

How Sparse Probability Maps Shape Mixture-of-Experts Routing

Tomás Brogueira, Marcos Treviso, Miguel Couceiro

arXiv 2610.06677首次发表:更新:

发表机构

Técnico, Universidade de Lisboa; INESC-ID; Instituto de Telecomunicações; ELLIS Unit Lisbon; Gandara AI(里斯本大学高等技术学院; INESC-ID研究所; 电信研究所; ELLIS里斯本分部; Gandara AI公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨稀疏概率映射(如sparsemax、entmax)在MoE路由中的训练行为,发现其稀疏性依赖于与学习得分的共同适应,虽不改善验证损失,但能显著降低推理时对专家数量的敏感性。

AI 中文摘要

混合专家(MoE)路由器通常对路由器得分应用softmax并保留前K个专家,使得每个令牌恰好使用K个专家。诸如sparsemax、alpha-entmax和normmax等稀疏诱导概率映射可以自适应地将精确的零分配给选定的专家,因此即使在使用相同的top-K机制时,它们似乎也能提供依赖于令牌的专家参与。在本研究中,我们探讨了这种稀疏性在训练过程中是否以及如何存续。我们使用softmax、1.5-entmax、sparsemax和2-normmax训练了匹配的300M和1B规模的top-2 MoE语言模型,并发现这些映射在训练后表现截然不同:在1B规模下,entmax丢弃的概率质量比softmax少30%,同时几乎从不丢弃选中的专家;sparsemax保留了最多的概率质量;而normmax将21%的令牌路由到单个专家。这些结果并非映射本身所固有。每个映射仅在两个最大得分之间的差距达到固定阈值时才丢弃选中的专家,而训练后的路由器在它们学习到的得分分布上有所不同:entmax路由器学习到的得分分布跨度大约是softmax的一半,这使其top-2差距保持在阈值以下,而sparsemax和normmax共享相同的阈值,却学习到不同的差距分布,因此专家参与也不同。因此,路由器会使其得分与映射共同适应,而映射产生零的能力本身并不能决定专家的参与。尽管没有任何稀疏映射在验证损失上优于softmax,但它们使训练后的模型在推理时对选择更多专家的敏感性大大降低:使用K=2训练的sparsemax在运行K=8时损失0.02 nats,而softmax损失0.58 nats。我们的结果表明,自适应MoE路由必须围绕概率映射和学习得分的联合行为来设计,而不是仅围绕映射本身。

英文摘要

Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, 1.5-entmax, sparsemax and 2-normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores reaches a fixed threshold, and the trained routers differ in the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. While none of the sparse maps improves validation loss over softmax, they make the trained models far less sensitive to selecting more experts at inference: sparsemax trained with K=2 loses 0.02 nats when run with K=8, where softmax loses 0.58. Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.

Comments22 pages, 6 figures, 9 tables. Under review at ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑