发表机构
Indian Institute of Technology Roorkee(鲁尔基印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对扩散语言模型中线性插值(LERP)不适用于超球面嵌入空间的问题,提出球形软掩码(S-SM)方法,在1.69亿参数MDLM上实现了MAUVE和困惑度的显著提升。
AI 中文摘要
软掩码可加速掩码扩散语言模型(MDLMs)的收敛,现有公式在原始嵌入空间中采用线性插值(LERP)实现该混合,隐含将该空间视为欧几里得空间。本文分析MDLMs的嵌入空间,发现掩码与预测标记的嵌入在整个训练过程中保持约73°的恒定夹角,且嵌入范数在词汇频率排名上基本平稳,表明该空间为超球面几何,而LERP是不适合的插值基元。本文提出球形软掩码(S-SM),这是一种可直接替换的方法,通过超球面上的Fréchet均值聚合前k个预测,再用球面线性插值(SLERP)将该均值与掩码方向混合,最后恢复原生掩码范数。本文在已发布的1.69亿参数MDLM检查点的持续预训练上,针对多种推理时间步预算评估S-SM:SLERP反馈避免了LERP反馈引发的训练退化,在不同采样预算下,相比普通MDLM基准,MAUVE指标提升最高达2倍,相比TopK/LERP提升27.5%-56.1%,同时生成困惑度持续降低(相比基准降低16.9%-19.6%),且输出熵和收敛性基本保持不变。
英文摘要
Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (LERP) in the raw embedding space, which implicitly treats that space as Euclidean. We analyze the embedding space of MDLMs and find that the mask and predicted-token embeddings maintain a near-constant angle of (\approx 73^\circ) throughout training, while embedding norms remain essentially flat across vocabulary-frequency rank. These indicate a hyperspherical geometry, for which LERP is the wrong interpolation primitive. We introduce Spherical Soft-Masking (S-SM), a drop-in replacement that aggregates the top-(k) predictions with a Fr'echet mean on the hypersphere and blends this mean with the mask direction using spherical linear interpolation (SLERP), then restores the native mask norm. We evaluate S-SM on continued pre-training of a released 169M-parameter MDLM checkpoint across a wide range of inference-time step budgets, SLERP feedback avoids the training degradation that LERP feedback induces and delivers MAUVE gains of up to 2x over the vanilla MDLM baseline and 27.5-56.1% over TopK/LERP at various sampling budgets, alongside consistently lower generative perplexity (16.9-19.6% over the baseline), while leaving output entropy and convergence essentially unchanged.
Comments15 pages