发表机构
University of Illinois Urbana-Champaign; AWS AI(伊利诺伊大学厄巴纳-香槟分校; 亚马逊云科技人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
APEX通过请求级专家选择和块级深度自适应,动态调整推测解码策略,在vLLM中实现最高5.24倍加速,并显著减少浪费令牌。
AI 中文摘要
推测解码通过在接受目标模型验证之前草拟多个令牌来减少大语言模型的推理延迟,但其有效性取决于提议机制和草稿深度。固定配置无法应对生成过程中可预测性、重复性和接受度的变化,因此更深的草稿可能会增加计算浪费而无法获得相应的加速。我们引入了APEX,一种学习型控制器,通过请求级专家选择和块级深度自适应来平衡解码速度和草稿令牌浪费。APEX-Router为每个请求在EAGLE-3、n-gram和草稿模型推测之间进行选择,而APEX-Depth利用因果解码信号和最近的验证器反馈在每个验证块调整草稿长度。APEX将接受的草稿长度建模为删失生存反馈,学习位置相关的拒绝风险、块执行成本以及一种平衡吞吐量、接受进度和浪费令牌的动作效用。这使得控制器能够在保留目标模型验证过程的同时自适应推测。我们将APEX集成到vLLM中,并使用Qwen3-8B在六个工作负载上进行评估,相对于自回归解码实现了高达5.24倍的加速。在总体评估中,APEX-S实现了4.27倍的加速,而APEX-B实现了3.27倍的加速,与固定n-gram推测(k=16)相比,浪费令牌百分比相对减少了41.0%,为平衡加速和草稿令牌利用提供了不同的操作点。
英文摘要
Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback. APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens. This allows the controller to adapt speculation while retaining the target model's verification procedure. We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding. Across the aggregate evaluation, APEX-S achieves 4.27X speedup, while APEX-B achieves 3.27X speedup with a 41.0% relative reduction in wasted-token percentage compared with fixed n-gram speculation at k=16, providing distinct operating points for balancing acceleration and draft-token utilization.