arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32448cs.AIcs.CL

ForkLeft:面向前缀对齐的自回归到扩散蒸馏的熵优先展开

ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation

Junming Liu, Jicheng Wang, Yifeng He, Hao Chen, Jianzhong Qi

首次发表
浏览论文内容

中文总结 AI 辅助

ForkLeft通过熵优先展开分离学生生成与教师监督,解决自回归教师与扩散学生间的前缀不匹配,使DLM在不放弃并行解码的前提下获得NTP式推理能力,并在多个基准上显著提升性能。

中文摘要 AI 辅助

自回归下一个词元预测(NTP)赋予了语言模型强大的推理能力,而扩散语言模型(DLM)则提供了灵活的词元顺序和并行生成能力。我们探究DLM能否在不放弃其原生生成过程的前提下,通过蒸馏获得NTP式的推理能力。然而,直接蒸馏面临一个根本性的不匹配:自回归教师模型基于左侧前缀进行预测,而DLM可以同时基于两侧的词元进行条件生成。我们提出了ForkLeft,一个蒸馏框架,通过将学生的展开过程与教师监督分离来解决这一不匹配。在训练期间,学生首先执行熵优先的展开,提交不确定的位置并暴露潜在的分叉。随后,我们固定由此产生的学生前缀,并在相同上下文下蒸馏一个NTP教师,以答案正确性决定监督来源。在推理时,学生回归其原生的置信度优先并行解码。使用Qwen3-30B-A3B-Base,ForkLeft在全部十个基准上提升了Efficient-DLM-4B的性能,将MATH500从72.60%提高到79.60%,并持续优于三种替代设计。其增益随教师强度提升而扩大,并仅需500次更新即可泛化到SDAR-4B。在匹配规模下,蒸馏后的4B和8B学生在七个基准上超过了已发表的SDAR-Chat和OPDLM模型,表明DLM可以在不牺牲原生并行生成能力的情况下学习NTP式推理。代码和数据集将在论文被接收后发布。

英文摘要

Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through distillation without giving up their native generation process. Direct distillation, however, faces a fundamental mismatch: an autoregressive teacher predicts from a left prefix, whereas a DLM can condition on tokens on both sides. We introduce ForkLeft, a distillation framework that resolves this mismatch by separating the student's rollout from teacher supervision. During training, the student first performs entropy-first rollouts that commit uncertain positions and expose potential forks. We then fix the resulting student prefix and distill an NTP teacher under the same context, with answer correctness determining the supervision source. At inference, the student returns to its native confidence-first parallel decoding. With Qwen3-30B-A3B-Base, ForkLeft improves Efficient-DLM-4B on all ten benchmarks, raising MATH500 from 72.60% to 79.60% and consistently outperforming three alternative designs. The gains scale with teacher strength and generalize to SDAR-4B with only $500$ updates. At matched scale, the distilled 4B and 8B students exceed the published SDAR-Chat and OPDLM models on seven benchmarks, showing that DLMs can learn NTP-style reasoning without sacrificing native parallel generation. Code and datasets will be released upon acceptance.

补充信息

↑