MDLMPE:面向掩码扩散语言模型的分布感知位置编码
MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
针对掩码扩散语言模型位置编码对动态标记可用性结构不敏感的问题,提出MDLMPE,通过编码标记可用性等实现分布感知,在LLaDA、DREAM等实验中性能优于传统方法。
中文摘要 AI 辅助
掩码扩散语言模型(MDLM)支持并行生成与双向上下文建模,但其位置上下文与自回归(AR)模型存在根本差异:自回归解码会生成连续前缀,而MDLM去噪则会产生已揭示与掩码标记的动态非连续配置。RoPE等传统位置编码可捕捉序列顺序与成对位移,但对这种不断变化的标记可用性结构不敏感。为解决此局限,我们提出MDLMPE,一种专为掩码扩散设计的位置编码。据我们所知,MDLMPE是首个使位置表示明确感知不断变化的已揭示/掩码配置的方法。它将标记可用性表示为二进制序列,应用距离感知高斯加权,并通过余弦基投影所得模式以获得分布感知位置特征。这些特征被添加到标记嵌入中,并通过轻量MLP映射为角度偏移,以调制标准RoPE相位。在LLaDA与DREAM上的大量实验表明,MDLMPE在监督微调、预训练、零样本评估及块扩散设置中,总体优于传统位置编码方法。进一步 ablation 实验显示,可用性状态、高斯局部性、频谱基与嵌入注入的完整组合可产生最强结果。这些结果确立了不断变化的标记可用性分布是掩码扩散语言模型的有用位置信号。
英文摘要
Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain insensitive to this evolving token-availability structure. To address this limitation, we propose MDLMPE, a positional encoding designed specifically for masked diffusion. To the best of our knowledge, MDLMPE is the first method to make positional representations explicitly aware of the changing revealed/masked configuration. It represents token availability as a binary sequence, applies distance-aware Gaussian weighting, and projects the resulting pattern through a cosine basis to obtain distribution-aware positional features. These features are added to token embeddings and mapped by a lightweight MLP to angular offsets that modulate the standard RoPE phases. Extensive experiments on LLaDA and DREAM demonstrate that MDLMPE generally outperforms conventional positional encoding methods across supervised fine-tuning, pretraining, zero-shot evaluation, and block-diffusion settings. Further ablations show that the complete combination of availability state, Gaussian locality, spectral basis, and embedding injection yields the strongest result. These results establish the evolving token-availability distribution as a useful positional signal for masked diffusion language models.