发表机构
University of Southern California; New York University; Columbia University; University of California, Berkeley; Stevens Institute of Technology; University of Notre Dame; Nanyang Technological University(南加州大学; 纽约大学; 哥伦比亚大学; 加州大学伯克利分校; 史蒂文斯理工学院; 圣母大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究离散掩码扩散语言模型中因果预训练与双向去噪协调问题,提出PreDiff-LM模型,通过混合注意力机制改进无条件困惑度等指标,在多方面优于先前扩散基线,明确与优化AR模型的差距,为预训练因果主干调整提供补充机制。
AI 中文摘要
离散掩码扩散语言模型支持双向生成和填充,但调整预训练的自回归(AR)变压器需要协调因果预训练和双向去噪。我们在注意力层面研究这个问题,而不是将AR权重重用本身视为新颖的。PreDiff-LM在观察到的提示中保留因果注意力,同时允许在掩码目标中进行完全双向注意力。在匹配的GPT-2 Medium、WikiText-103、90K步设置下,这种混合掩码将无条件困惑度从34.1提高到28.7,将MAUVE从0.71提高到0.78。注意力适应还与DiffuGPT风格的目标适应相结合,达到26.9的困惑度。预训练初始化将达到困惑度低于50所需的步数从约350K减少到8K。除了困惑度,PreDiff-LM还在重复、分布质量、四个零样本下游任务和人类偏好方面优于先前的扩散基线。结果表明混合注意力是调整预训练因果主干的一种补充机制,同时明确了与优化的AR模型在质量和推理效率上的差距。
英文摘要
Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requires reconciling causal pretraining with bidirectional denoising. We study this problem at the level of attention rather than claiming AR-weight reuse itself as novel. PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target. Under a matched GPT-2 Medium, WikiText-103, 90K-step setup, this hybrid mask improves unconditional perplexity from 34.1 to 28.7 and MAUVE from 0.71 to 0.78 over uniform bidirectional attention with the same AR initialization. Attention adaptation also composes with a DiffuGPT-style objective adaptation, reaching 26.9 perplexity. Pretrained initialization reduces the steps required to reach perplexity below 50 from about 350K to 8K, although a compute-matched fine-tuned AR model remains stronger at equal scale (18.9 versus 28.7). Beyond perplexity, PreDiff-LM improves repetition, distributional quality, four zero-shot downstream tasks, and human preference over prior diffusion baselines. The results position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized AR models.