arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于肘部法的MoE路由:一种用于专家选择的无需训练的推理时插件

Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

Robin Pan, Raymond Liu, Daniel Fang, Adelina Andrei, Rosa Wu

arXiv 2608.04401首次发表:更新:

发表机构

John A. Paulson School of Engineering and Applied Sciences, Harvard University; Department of Mathematics, Harvard University(哈佛大学约翰·A·保尔森工程与应用科学学院; 哈佛大学数学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无需训练的肘部法MoE推理时路由插件,动态调整每个token的专家数量,保持专家负载均衡,在六个基准测试中以5.3%的平均延迟降低量维持MoE模型准确率。

AI 中文摘要

混合专家(MoE)模型通过对每个token仅激活部分专家,在保持低推理计算量的同时实现模型扩展。然而,传统路由依赖固定的top-k选择,无论相关专家数量多少,模型都需消耗相同的计算量。我们提出肘部法路由,这是一种无需训练的推理时修改方法,可在每个token基础上动态调整专家数量。该方法检查排序后的路由概率分布,识别出分隔高、低概率专家的肘部点。我们发现大多数路由分布呈现适合该策略的清晰拐点,且通过理论和实验证明,肘部法路由可保持专家负载均衡。在最先进的MoE模型上进行的实验显示,其在六个基准测试中保持准确率的同时,平均延迟降低了5.3%。

英文摘要

Mixture-of-Experts (MoE) models enable model scaling while maintaining low inference-time compute by activating only a subset of experts per token. However, conventional routing relies on a fixed top-k selection, forcing the model to spend the same compute regardless of how many experts are relevant. We introduce elbow-based routing, a training-free inference-time modification that dynamically adjusts the number of experts on a per-token basis. Our method examines the sorted router probability distribution and identifies an elbow point that separates high- and low-probability experts. We find that most router distributions exhibit clear inflection points suitable for this strategy, and we show both theoretically and empirically that elbow-based routing preserves expert load balance. Experiments on a state-of-the-art MoE model demonstrate an average latency reduction of 5.3% while maintaining accuracy across six benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑