发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出A2D框架,无需训练即可将自回归模型的后训练权重更新应用于扩散语言模型,通过组合AR与扩散后训练更新,提升指令遵循、数学推理和编码能力。
AI 中文摘要
扩散语言模型(dLLMs)已成为自回归(AR)语言模型的一种有前景的替代方案,提供灵活的令牌更新顺序和并行解码。最近的dLLMs通常在扩散转换之前从预训练的AR模型初始化,以继承其学习到的表示。然而,在转换之后,它们通常忽略了其AR前身的广泛后训练生态系统。在这项工作中,我们表明这些现有的AR后训练权重更新可以有效地被回收利用以增强扩散模型。尽管AR到扩散转换带来了变化,直接将AR后训练权重更新添加到扩散基础模型仍然有效,使其性能接近通过直接扩散后训练所达到的水平。值得注意的是,AR和扩散后训练更新在权重空间中几乎是正交的,但在扩散模型中却引发了显著更一致的表征变化。它们不同的更新也是互补的:组合它们的权重可以保留两种机制带来的收益,并进一步改进后训练的扩散模型。基于这些发现,我们提出了A2D,一个简单的无需训练框架,利用现有的AR后训练资源来增强扩散模型。A2D可以将能力从AR后训练模型转移到扩散基础模型,并通过组合AR和扩散后训练更新来进一步改进已经后训练的扩散模型。在各种dLLMs中,包括Dream、DreamReasoner、DiffuCoder、Dream-Coder、Nemotron-Labs-Diffusion和DiffusionGemma,A2D可靠地改进了指令遵循、数学推理和编码,同时使用监督微调和强化学习更新,无需额外训练或推理时计算。
英文摘要
Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding. Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations. After the conversion, however, they typically ignore the extensive post-training ecosystem of their AR ancestors. In this work, we show that these existing AR post-training weight updates can instead be effectively recycled to enhance diffusion models. Despite the changes by AR-to-diffusion conversion, directly adding an AR post-training weight update to a diffusion base model remains effective, bringing its performance close to that achieved by direct diffusion post-training. Notably, AR and diffusion post-training updates are nearly orthogonal in weight space, yet induce substantially more aligned representation changes in the diffusion model. Their distinct updates are also complementary: composing their weights can retain gains from both regimes and further improve the post-trained diffusion model. Based on these findings, we propose A2D, a simple training-free framework for enhancing diffusion models with existing AR post-training resources. A2D can transfer capabilities from AR post-trained models to diffusion base models, and further improve already post-trained diffusion models by composing AR and diffusion post-training updates. Across various dLLMs, including Dream, DreamReasoner, DiffuCoder, Dream-Coder, Nemotron-Labs-Diffusion, and DiffusionGemma, A2D reliably improves instruction following, mathematical reasoning, and coding with both supervised fine-tuning and reinforcement learning updates, without additional training, or inference-time computation.