arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向块扩散语言模型的无偏在线蒸馏

Towards Unbiased On-Policy Distillation for Block Diffusion Language Models

Zaiquan Yang, Fei Wei, Yong Wang, Yudong Han, Yiyu Li, Zhuofan Zong, Gerhard Petrus Hancke, Xiangxiang Chu, Rynson WH Lau

arXiv 2610.05373首次发表:更新:

发表机构

City University of Hong Kong; Alibaba Group; Beijing Institute of Technology; The Chinese University of Hong Kong; City University of Hong Kong (Dongguan)(香港城市大学; 阿里巴巴集团; 北京理工大学; 香港中文大学; 香港城市大学(东莞))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对块扩散语言模型在线蒸馏中的上下文错位和过度自信崩溃问题,提出Un-OPD框架,通过边界过滤和置信度校准稳定训练,在数学与代码任务上性能更优且训练时间减半。

AI 中文摘要

在线蒸馏(OPD)已成为语言模型有效的训练后范式,近期的研究将其扩展到块扩散语言模型(BDLMs)。然而,现有研究几乎只关注小块大小,对将蒸馏用于更大块的学生模型的研究尚不充分。在本工作中,我们研究这一领域,并揭示了两种导致严重训练不稳定的关键优化偏差。首先,教师和学生之间的块边界不匹配导致上下文错位,提供扭曲的监督信号,误导学生的解码过程。其次,即使在上下文对齐的情况下,OPD中固有的优化偏差(学生倾向于快速吸收高支持信号,而在低支持更新上滞后)会驱动过早的置信度激增,使较弱的学生陷入灾难性的过度自信崩溃。为解决这些问题,我们提出Un-OPD,一个无偏的在线蒸馏框架,具有两个创新点以稳定BDLM训练。首先,Un-OPD引入边界感知的步骤过滤策略,消除上下文错位的解码步骤。其次,Un-OPD提出通过支持度再平衡的置信度校准来调节高支持位置的优化强度,从而避免过度自信崩溃。除了稳定性,我们还引入了一种回放重用机制以减少回放生成的开销。在数学推理和代码生成基准上的大量实验表明,Un-OPD能持续稳定训练并提供优越的性能,同时将墙钟训练时间减少约一半。

英文摘要

On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexplored. In this work, we investigate this regime and reveal two critical optimization biases that induce severe training instability. First, mismatched block boundaries between teacher and student cause \textbf{\textit{context misalignment}}, providing distorted supervisory signals that misguide student decoding. Second, even under aligned contexts, an \textbf{\textit{intrinsic optimization bias}} in OPD, where the student tends to rapidly absorb high-support signals while lagging on low-support updates, drives a premature confidence surge that traps weaker students in catastrophic overconfidence collapse. To resolve these, we propose \mbox{\textbf{Un-OPD}}, an unbiased on-policy distillation framework with two novelties for stabilizing BDLM training. First, Un-OPD introduces a boundary-aware step filtering strategy that eliminates context-misaligned decoding steps. Second, Un-OPD proposes moderating optimization intensity at high-support positions via a support-rebalanced confidence calibration, thereby bypassing overconfidence collapse. Beyond stability, we also introduce a rollout reuse mechanism to reduce rollout generation overhead. Extensive experiments on math reasoning and code generation benchmarks show that Un-OPD consistently stabilizes training and delivers superior performance while reducing wall-clock training time by approximately half.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑