发表机构
Institute for Artificial Intelligence, Peking University; School of Integrated Circuits, Peking University; College of Engineering, Peking University; School of Airspace Science and Engineering, Shandong University; Beijing Advanced Innovation Center for Integrated Circuits(北京大学人工智能研究院; 北京大学集成电路学院; 北京大学工学院; 山东大学空天科学与工程学院; 北京集成电路高精尖创新中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散-自回归推测解码中草稿与验证预算分配问题,提出互惠引导(RecGuide)框架,利用草稿与验证的互惠可预测性动态编排,在不同并发下提升吞吐量,最高达1.8倍加速。
AI 中文摘要
扩散式草稿生成结合自回归(AR)验证已成为高效推测解码的一种有前景的范式。以Nemotron-Labs-Diffusion为代表的近期自推测模型,通过在共享骨干网络内统一草稿生成与验证,进一步简化了推测流水线,同时实现了更长的接受长度。然而,聚合吞吐量与每请求吞吐量之间的帕累托前沿仍未得到充分探索。在低并发下,顺序的草稿-验证执行每轮需要两次模型前向传播,限制了每次前向的有效令牌数(TPF)。相比之下,在高并发下,更长的草稿会带来日益昂贵的计算开销,迫使单个请求在受限的推测预算下运行,从而无法充分利用全骨干草稿生成器。我们的关键观察表明,草稿生成与验证表现出互惠的可预测性。草稿逻辑可以预判可能的验证不匹配,而最近的验证结果则能预测未来的草稿效用及合适的块大小。基于这一观察,我们引入了互惠引导(RecGuide),一种运行时草稿-验证编排框架,使推测解码适应不同的服务负载。RecGuide在低并发下通过验证重叠的草稿生成利用空闲计算能力,同时在工作负载日益计算密集时,动态分配请求特定的草稿块大小。在广泛的并发水平上的实验表明,与原始自推测相比,吞吐量持续提升,实现了高达1.8倍的速度提升。
英文摘要
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to $1.8\times$ speedup.