arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知进退:均匀态扩散语言模型中的正确令牌保留

Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models

Mojtaba Nafez, James Henderson

arXiv 2610.01275首次发表:更新:

发表机构

EPFL; Idiap Research Institute(瑞士洛桑联邦理工学院; 伊迪亚普研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对均匀态扩散语言模型在自我纠正时过度修改正确令牌的问题,提出正确令牌保留正则化(CTR-Reg)辅助损失,在不改变采样器的情况下,显著提升干净令牌准确率并降低生成困惑度,同时保持多样性。

AI 中文摘要

均匀态扩散模型(USDMs)可以在任意去噪步骤中修改任意令牌,这使得它们能够纠正自身的错误,这是相对于掩码扩散的一个关键优势。然而,自我纠正既需要修改不正确的令牌,也需要保留正确的令牌,而我们表明当前的USDMs缺乏后者。即使在贪婪尾部解码下,最先进的USDMs(DUO、UDLM和均匀噪声SEDD)在每一步都会持续修改512个位置中的173至270个,而这些大规模、不协调的编辑会破坏样本多样性。一个随机令牌破坏实验将该缺陷追溯到模型本身:它们重建干净令牌和破坏令牌的准确率几乎相同,尽管干净令牌是更容易的目标。对验证集NELBO的分解表明,训练几乎不奖励保留:不正确的预测在破坏位置受到重罚,但在干净位置几乎不受惩罚。我们提出了正确令牌保留正则化(CTR-Reg),一种简单但有效的辅助损失,它训练模型保留未被前向过程扰动的令牌,并且不需要改变采样器。CTR-Reg在六个基准测试中平均将干净令牌准确率提高了26.5个百分点,同时几乎不改变破坏令牌的准确率,其每步修订收敛到仅3至11个位置。仅用五个贪婪尾部步骤,所有三种模型在CTR-Reg下的生成困惑度降低了一半以上,同时保持了多样性,并且这些收益在采样预算范围内持续存在。我们的结果确定了正确令牌保留是自我纠正扩散语言模型的一个关键缺失要素,并展示了一种有效的修复方法。

英文摘要

Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173--270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3--11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.

Comments38 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑