arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38385cs.AIcs.CL

微调扩散语言模型:上下文选择与目标加权

Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting

Loay Mualem, Lluís Pastor-Pérez, Vinh Tong, Andrei Manolache, Tanja Bien, Steffen Staab, Mathias Niepert

首次发表
浏览论文内容

中文总结 AI 辅助

GoldiMask通过子模目标优化上下文选择并对目标加权,提升离散扩散语言模型微调效果,在推理和代码生成任务上取得更高准确率。

中文摘要 AI 辅助

离散扩散语言模型的监督微调会掩蔽部分响应令牌,并训练模型从可见上下文中恢复其原始值。因此,掩蔽模式决定了模型可用的上下文以及模型学习预测的令牌。均匀随机掩蔽没有明确考虑这些选择之间的相互作用。我们提出了GoldiMask,它通过近似最大化一个子模目标来选择要作为上下文揭示的令牌。该目标利用模型信号来平衡揭示令牌的益处与其作为预测目标的价值。GoldiMask随后根据剩余目标从所选上下文中获益的程度及其剩余学习潜力对其进行加权。在三个骨干网络和三个训练数据集上,GoldiMask在大多数评估设置中取得了最高的平均准确率,在推理和代码生成方面均展现出提升。组件消融实验表明,上下文选择和目标加权都对性能提升有所贡献。此外,在置信度阈值并行解码下,GoldiMask在GSM8K和MATH-500上减少了解码迭代次数,同时在较高置信度阈值下保持了相当的准确率。

英文摘要

Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by approximately maximizing a submodular objective. This objective uses model signals to balance the benefit of revealing tokens against their value as prediction targets. GoldiMask then weights the remaining targets according to how they benefit from the selected context and their remaining learning potential. Across three backbones and three training datasets, GoldiMask achieves the highest average accuracy in most evaluated settings, demonstrating gains on both reasoning and code generation. Component ablations show that both context selection and target weighting contribute to the gains. GoldiMask also reduces decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding, while maintaining comparable accuracy at higher confidence thresholds.

补充信息

↑