arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02438cs.AI

xPress:推测解码中扩散草稿模型的并行优化

xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, Naigang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

xPress是恢复扩散草稿模型因果依赖的并行优化方法,在Qwen3-8B的7个基准上,可提升推测解码的接受长度与端到端吞吐量。

中文摘要 AI 辅助

诸如dFlash之类的块扩散草稿模型可在单次前向传播中生成一整块草稿token,大幅降低推测解码中多token草稿生成的开销。单次离散去噪过程的关键最终步骤是利用每个位置的logit分布,有条件地独立采样token,因此生成的草稿是各位置的边际分布,而非联合分布:无法保证任何草稿token依赖于其前序token。这种独立采样的边际分布易生成单个token看似合理,但在目标模型的条件验证下联合概率极低的序列,这会导致早期拒绝并限制接受长度。为解决该问题,我们提出xPress,用于恢复扩散草稿模型中缺失的因果关系。xPress是一种轻量级因果优化器,可通过并行优化一次性协调整个扩散块,在无需逐token循环的情况下恢复并传播草稿间的因果依赖。在Qwen3-8B模型上,针对7个数学、代码和聊天基准测试,与原始dFlash扩散草稿模型相比,xPress平均将接受长度提升约30%(最高提升56%),端到端解码吞吐量平均提升约1.3倍(最高提升1.7倍)。

英文摘要

Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative decoding. The crucial final step of the single-pass discrete denoising process involves using the logit distribution at each position to sample conditionally independent tokens. The resulting draft is thus a set of per-position marginals, rather than a joint distribution: no draft token is guaranteed to depend on its predecessors. Such independently sampled marginals tend to produce sequences with tokens that are individually likely, but jointly improbable under the target model's distribution, which verifies each token conditionally. This can cause early rejection and limits acceptance length. To address this, we propose xPress as a means to restore the missing causality in diffusion drafters. xPress is a lightweight causal refiner that reconciles the whole diffusion block at once through parallel refinement, restoring and propagating causal dependencies across the draft without a token-by-token loop. On Qwen3-8B, across seven math, code, and chat benchmarks, xPress raises acceptance length by about 30% on average (up to +56%) and its end-to-end decoding throughput by about 1.3 on average (up to 1.7) compared to the original dFlash diffusion drafter.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • IBM(国际商业机器公司(IBM))

机构由 AI 辅助整理,请以论文原文为准。

↑