arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26612cs.PF

启动受限且可替代:为何三种推理优化在混合专家(MoE)模型中无法取得成效

Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models

Gokulakannan Sakthivel, Jerry Wu, Amogh Rajendra, Giriprasad Radhakrishnan

AI总结:

本文在三个MoE模型上测试三种推理优化,发现其因模型等待内核启动、专家可替代等原因失效,且路由器精度与输出质量为可分离目标。

AI中文摘要:

混合专家(MoE)模型会将每个token路由至众多专家网络中的少数几个,且该路由过程具有数据依赖性,这是标准推理优化未考虑到的特性。本文在OLMoE-1B-7B、DeepSeek-V2-Lite和Qwen3-30B-A3B三个模型上,实际测量了三种推理优化的效果:融合Triton内核单独可实现5.6倍至9.0倍的加速,但端到端仅为0.999倍,而测量得到的上限为1.07倍,原因是模型将时间花费在每次前向传播时约一千次内核启动的等待上,而非内核所优化的算术运算上。INT4量化平均会改变每个token位置8个选定专家中的0.53个,但通过全精度权重重新执行这些改变后的路由,仅能复现2.7%的质量损失,这使得专家具有可替代性而非专业性。移除全部23个this URL图断点(现有研究视为结构性修复的前置步骤)会使模型速度降低三分之一。第四个结果将三者关联:将路由器保持在FP16精度可降低20%的漂移,同时提升损失,因此路由保真度和输出质量是可分离的目标。所有数值均来自已提交的每个token路由转储重新计算得到。

英文摘要:

Mixture-of-Experts (MoE) models route each token to a few of many expert networks, and that routing is data-dependent in a way standard inference optimizations do not expect. This paper measures what three of them actually deliver on OLMoE-1B-7B, DeepSeek-V2-Lite, and Qwen3-30B-A3B. Fused Triton kernels reach 5.6x to 9.0x in isolation but 0.999x end to end against a measured 1.07x ceiling, because the model spends its time waiting on roughly a thousand kernel launches per forward pass rather than on the arithmetic those kernels improve. INT4 quantization changes on average 0.53 of the eight selected experts per token position, yet replaying exactly those changed routes through full-precision weights reproduces only 2.7% of the quality loss, which makes the experts substitutable rather than specialized. Removing all 23 graph breaks from PyTorch's compiler, the step prior work treats as the structural fix, makes the model three times slower. A fourth result ties the three together: leaving the routers in FP16 lowers drift by 20% while raising loss, so routing fidelity and output quality are separable objectives. Every number recomputes from committed per-token route dumps.

↑