ASPIRE:面向长上下文LLM推理的异步批处理自推测解码
ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
AI总结:
ASPIRE提出非同步批处理自推测解码框架,通过混合前向、在线调度和起草内刷新,在长上下文推理中实现1.70-4.58倍吞吐量加速,较先前最优方法平均提升27%。
AI中文摘要:
长上下文LLM推理受注意力机制瓶颈限制,其重复的KV缓存读取使得解码过程受内存限制。自推测解码通过稀疏注意力起草令牌并用全注意力验证来缓解这一问题,但现有的批处理方法仍然保持同步:批次中的所有请求共享单一的起草-验证调度,尽管最优起草长度在不同请求间差异很大,并在每个请求内部动态变化。我们提出ASPIRE,一个基于三个组件的非同步批处理自推测解码框架。首先,统一的混合前向传播允许起草和验证请求在同一批前向传播中共存,消除了全局起草-验证阶段的需求。其次,轻量级在线推测调度器利用每请求接受率估计和批次感知成本模型,让每个请求独立选择何时验证。第三,起草内刷新层在起草过程中于单个指定层执行全注意力,在每个起草步骤更新稀疏上下文以减少起草期间的陈旧性。在三个模型和五个推理及长上下文基准上,ASPIRE相较于自回归基线实现了$1.70$-$4.58\times$的解码吞吐量加速,并相较于最强的先前自推测基线将平均加速比提高了约$27\%$。
英文摘要:
Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a single draft-verify schedule, even though the optimal draft length varies widely across requests and changes dynamically within each request. We propose ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components. First, a unified mixed forward allows drafting and verifying requests to coexist in the same batched forward pass, removing the need for global draft-verify phases. Second, a lightweight online speculation scheduler uses per-request acceptance-rate estimates and a batch-aware cost model to let each request independently choose when to verify. Third, an intra-draft refresh layer performs full attention at a single designated layer during drafting, updating the sparse context at every draft step to reduce staleness during drafting. Across three models and five reasoning and long-context benchmarks, ASPIRE achieves $1.70$-$4.58\times$ speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\%$ over the strongest prior self-speculative baselines.