AI 中文总结
研究针对大规模专家混合模型推理效率受专家激活模式影响的问题,提出成本感知推测解码框架EcoSpec,通过纳入预测的边际专家激活成本进行草稿选择,在多个模型和基准测试中减少活跃专家足迹、提高解码速度。
AI 中文摘要
稀疏专家混合(MoE)模型已成为扩展大语言模型(LLMs)的重要方法,但其推理效率很大程度上取决于专家激活模式。推测解码(SD)通过并行验证多个草稿令牌来加速自回归生成,然而现有草稿选择策略主要优化接受可能性。在大规模MoE模型中,选择草稿令牌还决定了验证期间激活的专家联合。我们观察到置信驱动的SD会引入“专家分散”:高概率草稿令牌可能会路由到不相交的专家,增加专家权重内存流量并降低推测带来的加速。基于此,我们在MoE推理的非均匀内存成本结构下重新审视草稿树选择。我们提出了EcoSpec,一个成本感知推测解码框架,将预测的边际专家激活成本纳入草稿选择。通过轻量级专家预测器和动态专家缓冲区,EcoSpec倾向于在重用当前验证集已覆盖的专家的同时保持高接受可能性的草稿路径,而不修改目标模型验证规则。我们在三个大规模MoE模型上评估了EcoSpec,包括DeepSeek-V3.1(671B)、Qwen3-235B-A22B和GPT-OSS-120B,并在推理、编码、问答和对话基准测试中进行了评估。EcoSpec始终减少活跃专家足迹并提高端到端解码速度,实现高达1.62倍的加速。这些结果表明,考虑专家激活成本对于大规模MoE模型中的高效推测解码很重要。
英文摘要
Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference efficiency depends strongly on expert activation patterns. Speculative decoding (SD) accelerates autoregressive generation by verifying multiple draft tokens in parallel, yet existing draft selection strategies primarily optimize acceptance likelihood. In large-scale MoE models, however, selecting draft tokens also determines the union of experts activated during verification. We observe that confidence-driven SD can introduce \textit{expert scattering}: high-probability draft tokens may route to disjoint experts, increasing expert-weight memory traffic and reducing the speedup from speculation. Motivated by this observation, we revisit draft-tree selection under the non-uniform memory-cost structure of MoE inference. We propose \textsc{EcoSpec}, a cost-aware speculative decoding framework that incorporates predicted marginal expert activation cost into draft selection. With a lightweight expert predictor and a dynamic expert buffer, \textsc{EcoSpec} favors draft paths that preserve high acceptance likelihood while reusing experts already covered by the current verification set, without modifying the target-model verification rule. We evaluate \textsc{EcoSpec} on three large-scale MoE models, including DeepSeek-V3.1 (671B), Qwen3-235B-A22B, and GPT-OSS-120B, across reasoning, coding, question-answering, and dialogue benchmarks. \textsc{EcoSpec} consistently reduces active expert footprints and improves end-to-end decoding speed, achieving up to $1.62\times$ speedup. These results show that accounting for expert activation cost is important for efficient speculative decoding in large-scale MoE models.