arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先共享,再路由剩余部分:一种用于令牌自适应混合专家(MoE)计算的统一框架

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

Gongli Zhang, Zhulin Liu, C. L. Philip Chen

arXiv 2608.10392首次发表:更新:

AI 中文总结

该研究针对MoE模型决策间的依赖问题,提出UniF-MoE统一框架,先共享再路由剩余部分,在DomainBed和GLUE实验中提升性能并降低计算资源消耗。

AI 中文摘要

混合专家(MoE)模型近期已不再局限于路由固定数量的完整专家。共享专家设计可保留可复用知识,细粒度方法在专家内部调整计算,动态路由器则适配活跃专家的数量。然而这些决策通常是独立做出的,忽略了一个基本依赖关系:提取可复用计算会改变剩余内容以及剩余部分所需的专家容量。我们通过将稀疏升级的前馈专家分解为键值通道来研究这一依赖关系。协同激活的专家在一部分值位置上对齐;移除这些位置会改变专家偏好;更大的共享覆盖范围与更低的剩余专家需求相关。这些观察引出了一条原则:先共享,再路由剩余部分。我们将其实例化为UniF-MoE,一种用于令牌自适应MoE计算的统一框架。每个专家被划分为对齐块。共享需求得分设置共享块数量和路径权重,键原型选择共享内容,互补需求通过累积路由质量确定剩余专家数量。Gram正则化器分离并归一化路由器嵌入,促进多样的路由方向、稀疏的专家重叠以及简单的路由几何。在DomainBed和GLUE上的实验表明,这种统一设计相较于代表性的静态和动态MoE,提升了预测性能,同时减少了活跃计算量、推理延迟和内存占用。代码可在this https URL获取。

英文摘要

Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers adapt the number of active experts. Yet these decisions are usually made independently, overlooking a basic dependency: extracting reusable computation changes both what remains and how much expert capacity the remainder needs. We study this dependency by decomposing sparsely upcycled feed-forward experts into key-value channels. Co-activated experts align at a subset of value positions; removing these positions changes expert preference; and greater shared coverage is associated with lower residual expert demand. These observations lead to one principle: share first, then route what remains. We instantiate it in UniF-MoE, a unified framework for token-adaptive MoE computation. Each expert is partitioned into aligned blocks. A shared-demand score sets the shared block count and pathway weight, key prototypes select the shared content, and the complementary demand determines the residual expert count through cumulative routing mass. A Gram regularizer separates and normalizes router embeddings, promoting diverse routing directions, sparse expert overlap, and a simple routing geometry. Experiments on DomainBed and GLUE show that this unified design improves predictive performance over representative static and dynamic MoEs while reducing activated computation, inference latency, and memory. Code is available at https://github.com/existence0420/UniF-MoE.

Comments11 pages, 6 figures, and 7 tables; includes supplementary material. Code is available at https://github.com/existence0420/UniF-MoE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑