发表机构
KAIST AI; University of Toronto; Vector Institute(韩国科学技术院人工智能学院; 多伦多大学; 向量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Pivot-SD通过信息增益选择高影响承诺进行自蒸馏,仅用200个问题和四次展开即提升LLaDA-8B-Instruct在数学和代码基准上的表现。
AI 中文摘要
掩码扩散语言模型(dLMs)为复杂推理提供了一种有前景的并行替代自回归模型的方法。然而,它们面临一个独特的信用分配挑战,因为在去噪过程中的少数几次承诺会急剧降低剩余掩码位置的不确定性,并塑造响应的大部分内容。大多数针对dLMs的后训练方法并未利用这一信号来决定训练哪些令牌:它们通常训练最终文本或将奖励分配给整个去噪步骤,而不是选择塑造响应的个别承诺。我们提出了Pivot-SD,一个高效的无监督离线自蒸馏框架,仅监督这些高影响的承诺(枢轴)。Pivot-SD使用信息增益度量来选择枢轴,该度量衡量剩余掩码位置的不确定性降低。来自成功轨迹的枢轴使用交叉熵训练,而来自失败轨迹的枢轴使用针对性的非似然训练,其余失败轨迹保持不变。仅使用200个问题和每个问题四次展开,Pivot-SD在数学和代码基准上优于全序列SFT和预算匹配的扩散RL基线,改进LLaDA-8B-Instruct。
英文摘要
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
CommentsEMNLP 2026 Main (Oral)