arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

减少扩散语言模型中的预训练-生成不匹配

Reducing Pretraining-Generation Mismatch in Diffusion Language Models

Xiaocheng Lu, Huabin Liu, Song Guo, Jianguo Li

arXiv 2608.09424首次发表:更新:

发表机构

Inclusion AI; The Hong Kong University of Science and Technology(Inclusion AI; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对扩散语言模型的预训练-生成不匹配问题,提出 PCD 方法,在不改变推理模式的情况下,在 LLaDA2-Mini 和 Qwen 模型上显著提升了性能。

AI 中文摘要

自回归语言模型在训练与使用阶段保持一致:生成以干净的提示词为条件,训练时则根据干净的左侧上下文预测后续 token。扩散语言模型提供并行去噪能力,但原生扩散语言模型(dLLM)预训练可能会随机损坏提示词和续段 token,削弱提示词条件生成所需的干净前缀接口。我们针对提示词续段识别出这种不匹配,并提出 PCD(前缀条件扩散),这是一种将自回归前缀监督与无偏移后缀去噪相结合的预训练目标。在预训练目标层面,PCD 在持续预训练中修改了注意力掩码、损坏掩码和标签构建;它不需要自回归解码器、验证器或新的推理模式。通过对干净前缀侧进行自回归监督,仅对未知续段应用扩散,PCD 使局部训练接口更接近块扩散模型在评估时的查询方式。我们进一步将样本内前缀条件与样本间目标混合分离,从而能够单独识别局部对齐信号,无需使用可选的批次级混合旋钮。在 LLaDA2-Mini 和 Qwen-1.7B 骨干模型上,PCD 始终优于同系列原生 dLLM 稳定基线,在 LLaDA2-Mini 的六项基准平均主指标上取得了 4.2% 的相对提升(+2.56 个百分点),在 Qwen 的主要机制比较中实现了 14.2% 的相对提升(+4.86 个百分点)。这些结果表明,使预训练上下文分布与提示词条件生成对齐,可在不改变推理方式的情况下,显著缩小 dLLM 的续段差距。

英文摘要

Diffusion language models (dLLMs) generate text through iterative denoising, allowing multiple tokens to be predicted in parallel. However, pretraining may mask tokens throughout a sequence, whereas prompt continuation conditions on an intact prefix. This difference remains in conversion pipelines that denoise entire sequences during the stable stage. We propose Prefix-Conditioned Diffusion (PCD), which samples a boundary, preserves the prefix, and denoises the suffix. The training recipe also applies autoregressive supervision to the prefix. We evaluate PCD in the stable stage of a warmup, stable, and decay conversion pipeline, with inference unchanged. Matched experiments across model families show improvements in reasoning and coding over native diffusion training. A matched continuation study further shows that the advantage persists after a shared decay stage. In a separate reconstruction diagnostic, the full PCD recipe's advantage over native diffusion training reverses as more evaluation prefix tokens are masked. Controlled experiments show lower suffix reconstruction loss with a clean training prefix, an intact evaluation prefix, and no autoregressive loss.

Comments12 pages, 9 figures, 4 tables. Revised manuscript with expanded experiments and analysis

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑