发表机构
ShanghaiTech University; Shanghai Engineering Research Center of Intelligent Vision and Imaging; Tencent(上海科技大学; 上海智能视觉与成像工程技术研究中心; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出DA-DLM,通过位置导向的有向无环图显式建模词元依赖,解决扩散语言模型中的连贯性问题,在语言建模、生成和摘要任务上优于Block Diffusion,并匹配自回归模型性能。
AI 中文摘要
扩散语言模型(DLMs)通过迭代去噪掩码序列来生成文本,在每一步独立预测多个词元。这种条件独立性忽略了词元间的依赖关系,降低了连贯性——这一问题与非自回归翻译(NAT)中的多模态问题类似。借鉴有向无环Transformer(DAT)通过有向无环图(DAG)解决NAT中该问题的方法,我们提出了DA-DLM,该模型通过位置导向的DAG设计,将基于DAG的依赖建模适配到DLM的迭代设置中。位置导向的DAG将节点组绑定到固定的输出位置,使得在较早步骤中固定的词元通过学得的转换来锚定邻近的预测,并随着去噪过程演变,在锚点累积时聚焦于剩余的不确定性。在语言建模、开放式生成和摘要任务上,DA-DLM始终优于Block Diffusion,尤其是在较少的去噪步骤下,并且在保持并行生成优势的同时,与自回归模型性能相当。我们的代码在此https URL公开可用。
英文摘要
Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This conditional independence discards inter-token dependencies and degrades coherence-an issue that parallels the multi-modality problem in Non-Autoregressive Translation (NAT). Drawing on the Directed Acyclic Transformer (DAT), which tackles this problem in NAT via a Directed Acyclic Graph (DAG), we propose DA-DLM, a model that adapts DAG-based dependency modeling to DLMs' iterative setting through a position-oriented DAG design. The position-oriented DAG binds node groups to fixed output positions so that tokens fixed in earlier steps anchor neighboring predictions via learned transitions, and evolves with denoising to focus on remaining uncertainty as anchors accumulate. On language modeling, open-ended generation, and summarization, DA-DLM consistently outperforms Block Diffusion, especially under fewer denoising steps, and matches autoregressive models while preserving the parallel generation advantage. Our code is publicly available at https://github.com/jipy0222/DA-DLM.