arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AcceptMoE:用于高效MoE投机解码的承诺权重自规模验证器专家集

AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding

Shuang Liang, Hao Mark Chen, Zhiwen Mo, Qianzhou Wang, Guoyu Li, Lingxiao Ma, Wayne Luk

arXiv 2608.02989首次发表:更新:

AI 中文总结

AcceptMoE 是一种结合目标路由器分数与离线承诺概率的验证器侧专家选择器,可自动调整合格专家数量,在降低流量的同时提升 MoE 投机解码吞吐量,仅轻微损失准确率。

AI 中文摘要

投机解码在一次目标模型前向传播中验证 draft token 树。然而对于混合专家(MoE)目标,并行验证会激活所有树节点所选专家的并集,即便只有一小部分节点达到接受的输出。token 数量、激活专家并集大小和专家权重流量是不同的成本度量:减少 token 工作量不一定成比例缩小专家并集,在卸载场景下,传输流量还取决于缓存驻留情况。我们提出 AcceptMoE,一种验证器侧的专家选择器,它结合目标路由器分数与离线估计的承诺概率,并自动调整每个验证块的合格专家数量,无需用户指定专家预算。在卸载场景下,AcceptMoE 根据缓存驻留情况而非预测自然路由并预取对应专家权重来确定专家资格。尽管约束目标专家资格会改变模型分布,但在涵盖 3 个 MoE 目标和 4 个基准的 12 个模型-任务对上,AcceptMoE 的平均准确率比采用自然路由的 EAGLE-3 投机解码低 0.27 个百分点。在批大小为 1 时与 SGLang 一同部署,它在所有专家权重均在 GPU 内存中时达到基线吞吐量的 1.290 倍,在物理专家卸载下达到 2.06 倍,同时将主机到设备的流量降低 73.6% 至 77.1%。

英文摘要

Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verification can activate the union of the experts selected by all tree nodes, even though only a small subset of those nodes reaches the accepted output. Token count, activated-expert union size, and expert-weight traffic are therefore distinct cost measures: reducing the token workload need not shrink the expert union proportionally, and under offloading, transfer traffic also depends on cache residency. We introduce AcceptMoE, a verifier-side expert selector that combines target-router scores with offline-estimated commitment probabilities and automatically adjusts the number of eligible experts for each verification block, eliminating the need for a user-specified expert budget. Under offloading, AcceptMoE conditions expert eligibility on cache residency instead of predicting natural routes and prefetching the corresponding expert weights. Although constraining target-expert eligibility changes the model distribution, across 12 model-task pairs spanning three MoE targets and four benchmarks, AcceptMoE's mean accuracy is 0.27 percentage points lower than that of EAGLE-3 speculative decoding with natural routing. Served with SGLang at batch size one, it reaches 1.290 times the throughput of this baseline with all expert weights in GPU memory, and 2.06 times under physical expert offloading, while reducing host-to-device traffic by 73.6 percent to 77.1 percent.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑