arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33301cs.CLcs.AI

犹豫感知的扩散语言模型在线策略蒸馏

Hesitation-Aware On-Policy Distillation for Diffusion Language Models

Jianguo Huang, Lipeng Wan, Yanchen Deng, Bo An

AI总结:

针对扩散语言模型蒸馏中未提交提案信号被忽略的问题,提出犹豫感知在线策略蒸馏(HOPD),扩展教师分布匹配至所有掩码位置并利用后见之明加权,在数学和编码基准上取得最佳平均分并加速解码。

AI中文摘要:

扩散大语言模型(dLLM)通过迭代去掩码生成文本。在每个去噪步骤中,dLLM 在每个掩码位置提出一个词元,但解码器仅提交这些提案中置信度较高的子集。基于轨迹的在线策略蒸馏(TOPD)通过将学生模型与更强的教师模型匹配来构建此过程,但仅在已提交的位置进行匹配。我们认为这丢弃了大量有用信号,这些信号存在于未提交的提案中,即学生模型已做出预测但尚未足够自信去提交的位置。我们将这些提案称为犹豫。在我们对 SDAR-4B 学生模型的初步研究中,犹豫仅占可监督状态-位置对的 24%,却承载了 66% 的教师-学生差异。为利用这一信号,我们提出犹豫感知的在线策略蒸馏(HOPD),将教师分布匹配扩展到每个去噪步骤的每个掩码位置。由于犹豫的信息量并不相同,我们进一步利用完整轨迹的后见之明分配监督,对提案后来与最终词元不一致的位置以及第一步提案很少存活的块赋予更多权重。由于两个模型已在所有掩码位置产生分布,HOPD 相比 TOPD 不需要额外的前向传播。唯一额外成本是在更多位置评估损失。使用从 TraDo-8B-Instruct 蒸馏的 SDAR-1.7B 和 SDAR-4B 学生模型,HOPD 在五个数学和编码基准上,在静态和动态解码以及两种规模下均取得评估方法中的最佳平均分数。它还加速了解码。在 SDAR-4B 上,HOPD 学生模型比 TOPD 犹豫更少,每个去噪步骤多提交 11% 的词元,同时达到更高的准确率。

英文摘要:

Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the student to a stronger teacher, yet only at the committed positions. We argue that this discards much of the useful signal, which resides in the uncommitted proposals, where the student has made a prediction but is not yet confident enough to commit it. We call these proposals hesitations. In our pilot study on an SDAR-4B student, hesitations make up only 24% of supervisable state-position pairs but carry 66% of the teacher-student divergence. To exploit this signal, we propose Hesitation-Aware On-Policy Distillation (HOPD), which extends teacher distribution matching to every masked position of each denoising step. Because hesitations are not equally informative, we further allocate supervision using hindsight from the completed trajectory, placing more weight on positions whose proposal was later disagreed with the final token and on blocks where first-step proposals rarely survive. Since both models already produce distributions at all masked positions, HOPD requires no additional forward passes over TOPD. The only extra cost is evaluating the loss at more positions. With SDAR-1.7B and SDAR-4B students distilled from TraDo-8B-Instruct, HOPD achieves the best average score among the evaluated methods on five math and coding benchmarks, under both static and dynamic decoding and at both scales. It also speeds up decoding. On SDAR-4B, the HOPD student hesitates less and commits 11% more tokens per denoising step than TOPD, while reaching higher accuracy.

↑