发表机构
Tsinghua University; Beijing Institute of Mathematical Sciences and Applications; Wuhan University; MathonAI(清华大学; 北京数学科学与应用研究院; 武汉大学; 未知(暂未找到合适中文翻译))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在将预训练的自回归语言模型适配到均匀噪声扩散。通过建立不同模型间联系并推导转换,提出UNIFUSION方法。经实验评估,该方法能改善生成困惑度与单字熵权衡,在相同规模模型中多项指标表现优异。
AI 中文摘要
现有方法主要将预训练的自回归(AR)语言模型适配到掩码扩散,而本文直接将其适配到均匀噪声扩散,即采样时每个token仍可编辑。但跨损坏内核适配AR检查点具有挑战性,因现有DLMs使用不同目标和预测参数化。本文通过将SEDD、MDLM/GIDD、M2S和Neural CTMC的条件损失表示为模型反向速率上的单个广义Kullback-Leibler目标来建立联系,还推导了从干净token预测到具体分数、后验均值和退出速率/跳跃参数化的转换,提出了UNIFUSION方法。通过对124M和355M参数模型的系统评估,表明随着采样预算从16步增加到256步,UNIFUSION稳步改善了生成困惑度和单字熵之间的权衡,在相同规模的模型中,该方法在生成困惑度、单字熵以及WinoGrande、SIQA和BBH准确性方面表现出色。
英文摘要
Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across corruption kernels remains challenging because existing DLMs use different objectives and prediction parameterizations. We establish connections among SEDD, MDLM/GIDD, M2S, and Neural CTMC by expressing their conditional losses as a single generalized Kullback--Leibler objective over model reverse rates. We further derive conversions from clean-token predictions to concrete-score, posterior-mean, and exit-rate/jump parameterizations, yielding a shared \(x_0\) interface that supports switching between mask and uniform kernels. Building on these connections, we propose \ours{}, a simple continual pre-training approach for directly adapting pretrained GPT2 checkpoints to uniform-noise diffusion. Through systematic evaluation of 124M- and 355M-parameter models, we show that \ours{} steadily improves the trade-off between generative perplexity (GenPPL) and unigram entropy as the sampling budget increases from 16 to 256 steps. At 256 steps, \ours{}-S and \ours{}-M achieve GenPPL/entropy pairs of \(97.783/5.2626\) and \(71.516/5.6669\), respectively; no evaluated model at the same scale simultaneously outperforms \ours{} on both metrics. At both scales, \ours{} also achieves the highest WinoGrande, SIQA, and BBH accuracy among the compared diffusion models.