arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11742cs.CL

波纹枢轴搜索:扩散大语言模型的主动并行解码

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

  • Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学合作中位数网络创新中心)
  • Alibaba Group TMLR Group(阿里巴巴集团TMLR团队)
  • Hong Kong Baptist University(香港浸会大学)
  • A*STAR CFAR and Nanyang Technological University(新加坡科研局计算科学中心和南洋理工大学)
  • School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

Yushi Ye, Xu Chen, Haoyun Jiang, Jinsong Lan, Haihong Tang, Xiangtao Li, Mingming Gong, Ivor Tsang, Yanfeng Wang, Jiangchao Yao

AI总结:

提出RPS方法,利用dLLM解码的波纹效应,在保持生成质量的同时,大幅提升解码速度与准确率,适用于推理及代码生成任务。

AI中文摘要:

扩散大语言模型(dLLMs)已成为自回归语言模型的有力替代方案,具备通过并行解码实现大幅加快推理速度的潜力。现有并行解码调度器通常仅在每个位置满足对应标准后才确定该位置的内容,却忽略了提前确定部分位置对后续解码的益处。我们在dLLM解码中发现了一种波纹效应:主动确定一个中等熵的枢轴位置,可显著降低其余掩码位置的不确定性。这种不确定性降低能让后续步骤并行解封更多token,从而加快整体解码过程。为利用该波纹效应,我们提出了一种无需训练的新型解码方法——波纹枢轴搜索(RPS),该方法会寻找中等熵位置作为有前景的候选枢轴(即确定解码位置),并通过前瞻评估确定能产生最大下游收益的token分配方案(即确定解码内容)。在3种dLLMs和4个推理与代码生成基准上,RPS相较于标准解码器实现了4-10倍的挂钟速度提升,同时保持了生成质量;在多数设置下,其准确率较此前的前瞻基线最高提升5.49%,且吞吐量更高。当与KV缓存结合时,RPS相较于标准解码器进一步实现了最高18倍的挂钟速度提升。

英文摘要:

Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parallel decoding. Existing parallel decoding schedulers typically commit positions only after they meet a per-position criterion, overlooking how early commitments may benefit subsequent decoding. We identify a ripple effect in dLLM decoding: proactively committing a mid-entropy pivot position can induce a pronounced reduction in uncertainty across the remaining masked positions. This uncertainty reduction allows subsequent steps to unmask more tokens in parallel, thereby accelerating the overall decoding process. To exploit the ripple effect, we propose Ripple-Pivot Search (RPS), a novel training-free decoding method that seeks mid-entropy positions as promising candidate pivots (where to decode), and determines their token assignment that yields the greatest downstream benefit via lookahead evaluation (what to decode). Across 3 dLLMs and 4 reasoning and code-generation benchmarks, RPS achieves 4-10$\times$ wall-clock speedup over the standard decoder while preserving generation quality, and improves accuracy over the previous lookahead baseline by up to 5.49% while delivering higher throughput in most settings. When integrated with KV caching, RPS further achieves up to 18$\times$ wall-clock speedup over the standard decoder.

↑