重新思考扩散语言模型中用于并行解码的软令牌
Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对扩散语言模型并行解码中软令牌反馈机制未被系统研究的问题,提出无训练、几何感知的软令牌构建方法,并在四个预训练模型和四个基准上验证其优于标准并行解码。
AI中文摘要:
扩散语言模型(DLMs)通过在每次去噪步骤中预测并提交多个令牌来实现并行生成,然而它们可能生成单独合理但相互不一致的令牌。近期研究表明,软令牌可以通过使用先前解码步骤中模型预测分布构建的连续嵌入来表示不确定位置,从而缓解这一问题。然而,尽管软令牌通常被理解为保留预测不确定性,但软令牌反馈如何改善并行解码尚未被系统性地检验。在本文中,我们在冻结的预训练DLMs中研究这一问题,以在不增加额外训练影响的情况下考察软令牌反馈。为了在无训练设置中构建软令牌输入,我们识别出传统软令牌构建与预训练嵌入空间之间的几何不匹配。基于这一观察,我们提出了一种无训练、几何感知的软令牌构建方法。我们对软令牌反馈的分析表明,仅保留不确定性并不能完全解释其如何重塑后续预测。为了更好地解释软令牌反馈如何改善并行解码,我们提供了经验证据,表明它倾向于连贯的令牌序列。在四个预训练DLMs和四个数学与代码基准上,我们的方法优于标准并行解码和无训练的欧几里得软令牌基线。代码:此https URL。
英文摘要:
Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate individually plausible but mutually inconsistent tokens. Recent work shows that \emph{soft tokens} can mitigate this issue by representing uncertain positions with continuous embeddings built from the model's predictive distribution at the previous decoding step. However, although soft tokens are commonly understood as preserving predictive uncertainty, how soft-token feedback improves parallel decoding has not been systematically examined. In this paper, we investigate this question in frozen pretrained DLMs to examine soft-token feedback without the effects of additional training. To construct soft-token inputs in a training-free setting, we identify a geometric mismatch between conventional soft-token construction and the pretrained embedding space. Based on this observation, we propose a training-free, geometry-aware construction of soft tokens. Our analysis of soft-token feedback suggests that uncertainty preservation alone does not fully explain how it reshapes subsequent predictions. To better explain how soft-token feedback improves parallel decoding, we provide empirical evidence that it favors coherent token sequences. Across four pretrained DLMs and four math and code benchmarks, our method outperforms standard parallel decoding and a training-free Euclidean soft-token baseline. Code: https://github.com/kodaikawamura/rethinking-soft-tokens