投机极限:界定混合专家模型中的投机解码
The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts
浏览论文内容
中文总结 AI 辅助
本研究将MoE投机解码的预算选择建模为离线随机最短路径问题,通过诊断Oracle发现拒绝候选呈线性边界,揭示了平衡边际成本与预期进展的必要条件,为自适应在线启发式设计提供数学基础。
中文摘要 AI 辅助
混合专家(MoE)模型中的投机解码面临由输入相关专家加载引起的不稳定验证成本问题。为研究该过程的物理机制,我们将投机预算选择表述为参考序列上的离线随机最短路径(SSP)问题,并构建了一个诊断性Oracle,利用反事实模拟来核算MoE验证成本。对Qwen3-Coder与EAGLE-3配对在边际增量空间(Delta Space)中Oracle决策的详细分析表明,被拒绝的候选形成严格的线性边界。该结果证明,复杂的全局优化在局部受制于平衡边际成本与预期进展的必要条件($\frac{\Delta \mathbb{E}[Cost]}{\Delta \mathbb{E}[a]}$),为设计未来自适应在线启发式算法提供了严谨的数学参考点。
英文摘要
Speculative decoding in Mixture-of-Experts (MoE) models faces the problem of unstable verification cost caused by input-dependent expert loading. To study the physics of this process, we formulate speculation-budget selection as an offline Stochastic Shortest Path (SSP) problem over reference sequences and build a diagnostic Oracle that uses counterfactual simulation to account for MoE verification cost. A detailed analysis of the Oracle's decisions on the Qwen3-Coder and EAGLE-3 pairing, in the space of marginal deltas (Delta Space), shows that rejected candidates form a strict linear boundary. This result demonstrates that a complex global optimization is locally governed by a necessary condition balancing marginal cost against expected progress ($\frac{Δ\mathbb{E}[Cost]}{Δ\mathbb{E}[a]}$), providing a rigorous mathematical reference point for designing future adaptive online heuristics.