发表机构
Yandex Research; HSE University; Applied AI Institute; AXXX; T-Tech; Constructor University(Yandex研究院; 高等经济大学; 应用人工智能研究所; AXXX; T-Tech; 康斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种基于表示空间MMD的后训练方法,通过最小化生成与参考分布间的MMD,无需完整采样轨迹即可高效优化扩散语言模型,在困惑度、准确率和解码并行性上取得改进。
AI 中文摘要
我们提出了一种针对扩散语言模型(DLMs)的后训练方法,该方法在冻结的预训练DLM的特征空间中,最小化生成分布与参考分布之间的最大均值差异(MMD)。为了估计MMD,我们保留每个词元位置上的上下文特征,从而在单次提取器遍历中从每个序列获得多个观测值。我们通过策略梯度优化离散模型的目标,通过生成潜变量的直接微分优化连续模型的目标。在这两种情况下,直接从这些特征计算损失,无需完整的采样轨迹或联合训练的辅助模型,即可实现高效的后训练。实验表明,在OpenWebText上,以相当的熵获得了更低的生成困惑度,并在GSM8K上取得了更好的准确率-计算权衡。在具有混合掩码-均匀扩散的16B DMax-LLaDA2.0模型上,我们在数学和代码基准上以相似或更高的准确率提高了解码并行性。
英文摘要
We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.
CommentsTech Report. Code: https://github.com/yandex-research/mmd-dlm