arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于轨迹的策略蒸馏用于掩码扩散语言模型

Trace-Based On-Policy Distillation for Masked Diffusion Language Models

Haolin Ren, Ziyang Huang, Chenhao Yuan, Jun Zhao, Kang Liu

arXiv 2607.16872首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences(中国科学院自动化研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散大语言模型推理训练后处理难的问题,提出基于轨迹的策略蒸馏(TOPD)框架,通过在目标模型去噪轨迹上监督,利用教师模型获取令牌分布并以反向KL目标更新,在数学推理基准上提升了模型准确率且减少计算量。

AI 中文摘要

扩散大语言模型(dLLMs)是自回归生成的有前途的替代方案。然而,针对dLLMs的推理导向的训练后处理仍然具有挑战性。本文提出了基于轨迹的策略蒸馏(TOPD),这是一个教师监督框架,可在无需奖励估计的情况下将推理能力转移到目标dLLM。关键思想是在其自身的去噪轨迹上监督dLLM,关注形成最终响应的与轨迹对齐的令牌决策。具体来说,TOPD从目标dLLM中采样策略扩散轨迹,从教师模型在相应的部分去噪状态下获得教师令牌分布,并使用令牌级反向库尔贝克-莱布勒(Reverse-KL)目标更新目标dLLM。在数学推理基准上,TOPD使SDAR-4B-Chat在静态评估下提高了5.7,在动态评估下提高了4.5,匹配了其基于强化学习训练的对应模型TraDo-4B-Instruct的MATH500准确率,且推出轮次减少4倍。

英文摘要

Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relies on sparse rewards or value modeling. This paper proposes \textbf{trace-based on-policy distillation (TOPD)}, a teacher-supervised framework that transfers reasoning ability to a target dLLM without reward estimation. The key idea is to supervise a dLLM on its own denoising trajectory, focusing on the trace-aligned token decisions that form the final response. Specifically, TOPD samples on-policy diffusion trajectories from the target dLLM, obtains teacher token distributions from a teacher model on the corresponding partially denoised states, and updates the target dLLM with a token-level Reverse Kullback-Leibler (Reverse-KL) objective. This design preserves dense teacher supervision while aligning training with the model's own denoising states. On mathematical reasoning benchmarks, TOPD enables SDAR-4B-Chat to match the MATH500 accuracy of its RL-trained counterpart TraDo-4B-Instruct, with gains of +5.7 under static evaluation and +4.5 under dynamic evaluation. Compared with the RL-trained counterpart, TOPD achieves this with 4$\times$ fewer rollout rounds, corresponding to an estimated 96.0$\times$ to-accuracy model-compute speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑