发表机构
EPFL; Georgia Institute of Technology; Apple(洛桑联邦理工学院; 佐治亚理工学院; 苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PUMBA框架,通过轨迹感知训练、步骤间信息传递和时间反向传播,统一掩码扩散语言模型的训练与推理,提升性能并减少函数评估次数。
AI 中文摘要
掩码扩散模型(MDMs)通过每步解码多个标记来生成文本,但其训练与采样条件不同。模型在随机掩码序列上训练,而推理则遵循由模型自身预测塑造的轨迹。此外,每一步无法访问前一步的计算结果。近期方法从不同角度分别缓解这些局限,但这些选择之间的相互作用仍待探索。我们提出PUMBA,一个统一的轨迹感知训练框架,它在策略诱导轨迹的连续步骤上训练去噪器,在步骤间传递信息,并通过时间反向传播联合优化。对该设计空间的受控研究表明:i) 精确的训练-推理对齐因局部过拟合而失败,而较宽松的对齐仍能使训练掩码更接近推理时所见;ii) 通过每步的承诺,传递连续信息优于离散梯度估计器;iii) 随着时间反向传播跨越更多步骤,性能提升,我们对此提供了理论支持。结合这些组件,性能与同规模自回归模型的最佳检查点相当。基于这些发现,我们将PUMBA扩展到LLaDA-8B的监督微调,在整幅画布和块扩散生成中均改善了性能与函数评估次数(NFEs)之间的权衡。在性能匹配时,整幅画布生成中,它比使用双倍预算的标准微调最多减少22%的NFEs;在块扩散中,相同步数下比标准微调最多减少26%的NFEs。
英文摘要
Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on consecutive steps of policy-induced trajectories, passes information between steps, and optimizes them jointly by backpropagation through time. A controlled study of this design space shows that i) exact train--inference alignment fails due to local overfitting, whereas a looser alignment still brings training masks closer to those seen at inference; ii) passing continuous information outperforms discrete gradient estimators through the commitment at each step; and iii) performance improves as backpropagation through time spans more steps, which we support theoretically. Combined, these components match the best checkpoint of a same-size autoregressive model. Building on these findings, we scale PUMBA to supervised fine-tuning of LLaDA-8B, where it improves the trade-off between performance and number of function evaluations (NFEs) in both full-canvas and block diffusion generation. At matched performance, it needs up to 22% fewer NFEs than standard fine-tuning with twice the budget in full-canvas generation, and up to 26% fewer than standard fine-tuning for the same number of steps in block diffusion.