发表机构
University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过适配Qwen3.5混合架构(0.8B至9B)构建dQwen3.5扩散语言模型,发现混合骨干网络训练效率更高(节省约一半token),且解码性能与全注意力模型相当。
AI 中文摘要
适配预训练的自回归(AR)模型是构建扩散语言模型(DLM)的一种经济高效的途径。尽管几乎所有此类适配都始于全注意力Transformer,但AR建模已转向交错使用注意力层和RNN层的混合架构。这为适配带来了障碍:与注意力不同,RNN在结构上是因果的,且难以双向化。尽管存在这种不匹配,我们仍研究了此类骨干网络能否通过适配Qwen3.5在0.8B、2B、4B和9B规模上成为有效的DLM,从而产生了dQwen3.5系列。我们发现,混合骨干网络可以成为适配的高效起点:与全注意力对照组相比,混合模型达到给定训练损失所需的token数约少一半。在各规模下,dQwen3.5在任意顺序解码行为上与全注意力DLM相似,并在并行解码下表现强劲。
英文摘要
Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.