发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对均匀离散扩散模型提出无需训练的采样器 CRS,通过将所选 token 作为持久上下文插入后续输入,在 8-64 NFE 预算下实现更优的 GenPPL-熵权衡。
AI 中文摘要
均匀状态离散扩散模型会并行更新所有 token,同时保持每个位置可修订。即使常用的 top-p 规则在某个位置仅留下一个候选,该选择也仅影响当前的反向步骤,可在下次采样步骤修订。我们探究当所选假设变为后续预测的持久上下文时会发生什么,因此提出无需训练的采样器 committed reveal sampling(CRS),该采样器存储所选的 argmax token 并将其插入后续模型输入。我们的分析为延迟选择和保持所选 token 可见性提供了依据:在精确前向过程下,选择干净 token 的贝叶斯误差不会随噪声降低而增大;在简单的潜在模式模型中,保持所选 token 可见性有助于后续并行预测在序列级选择上达成一致。实证上,我们在 Duo-distilled 上进行配对实验,以将这种持久效应与单步 top-p 限制及标量温度缩放分离。在相同最终规则下,无 top-p 截断的 CRS 在 8 至 64 次函数评估(NFE)预算下,生成困惑度(GenPPL)低于固定 p=0.95 和 p=0.9 的基线。在 64 NFE 时,匹配单字熵的比较也显示 CRS 的 GenPPL 更低,形成更优的 GenPPL-熵权衡。Base Duo 在描述性比较中呈现相同趋势,而其他多样性和延续性指标可能对这些操作点有不同排序。这些结果表明,支持限制和持久上下文是控制该权衡的两种不同机制。
英文摘要
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. Even when the commonly used top-$p$ rule leaves only one candidate at a position, that choice affects only the current reverse step and can be revised at the next sampling step. We ask what changes when selected hypotheses instead become persistent context for later predictions. We therefore propose committed reveal sampling (CRS), a training-free sampler that stores selected argmax tokens and inserts them into subsequent model inputs. Our analysis gives a rationale for selecting later and for keeping selected tokens visible. Under the exact forward process, the Bayes error of selecting a clean token cannot increase as noise decreases, while in a simple latent-mode model, keeping the selected token visible helps later parallel predictions agree on the same sequence-level choice. Empirically, paired experiments on Duo-distilled then separate this persistent effect from single-step top-$p$ restriction and scalar temperature scaling. Under the same finalization rule, CRS without top-$p$ truncation reaches lower generative perplexity (GenPPL) than fixed $p=0.95$ and $p=0.9$ baselines across budgets of 8--64 function evaluations (NFE). At 64 NFE, the comparison at matched unigram entropy also gives lower GenPPL for CRS, yielding a more favorable GenPPL--entropy tradeoff. Base Duo shows the same direction in a descriptive comparison, while other diversity and continuation metrics can rank these operating points differently. These results identify support restriction and persistent context as distinct controls of that tradeoff.
Comments23 pages