arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35362cs.LGcs.AI

d-OPD:面向块扩散语言模型的未来感知在线策略蒸馏

d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models

Ruitao Liu, Qinghao Hu, Song Han

首次发表
浏览论文内容

中文总结 AI 辅助

针对块扩散语言模型蒸馏中教师与学生信息不对齐的问题,提出未来感知在线策略蒸馏方法d-OPD,通过纳入块内可见未来信息校正教师分布,提升Qwen3模型多基准平均分达4.0点并显著减少训练时间。

中文摘要 AI 辅助

大型语言模型(LLMs)通常以自回归(AR)方式生成文本,即一次预测一个词元。块扩散语言模型(dLLMs)则按顺序生成块,同时在每个块内并行去噪多个词元,为加速生成提供了一种有前景的方式。近期工作并非从头训练此类模型,而是通过蒸馏将强大的预训练AR模型适配为块dLLM。在线策略蒸馏(OPD)因其在学生当前策略生成的状态上监督学生,而非仅依赖固定的离线轨迹,而被广泛用于LLM训练。通过训练学生实际访问的状态,它减少了训练与生成之间的不匹配,并随着学生的发展提供更相关的监督。近期工作已将此思想扩展到AR到块扩散的转换。然而,此设置引入了监督中的根本性不匹配:在相同的训练状态下,块扩散学生和因果AR教师基于不同信息进行条件化。学生从整个部分去噪的块(包括可见的未来上下文)进行预测,而标准AR教师目标仅由因果前缀定义。因此,用于蒸馏的教师分布与学生可见的信息不完全对齐。为此,我们提出d-OPD,一种未来感知的在线策略蒸馏方法,通过在每个块内纳入可见的未来信息来校正AR教师分布,使其更好地与学生可见状态对齐,提供与学生所用信息更匹配的监督。在Qwen3模型从0.6B到8B的规模上,d-OPD在六个基准上的平均得分比OPDLM最多提升4.0分,并将训练时间减少1.35至1.58倍。代码可在https URL获取。

英文摘要

Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to $4.0$ points over OPDLM and reduces training time by $1.35$-$1.58\times$. The code is available at https://github.com/mit-han-lab/d-OPD.

发表机构

  • Tsinghua University(清华大学)
  • MIT(麻省理工学院)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

↑