arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

蒸馏后的连续扩散语言模型可少步生成代码——甚至仅需一步

Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang

arXiv 2609.04531首次发表:更新:

发表机构

Duke University; Tsinghua University(杜克大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出0.7B参数连续扩散代码生成模型PlaidQ,通过蒸馏实现少步或单步代码生成,16步模型性能超512步采样的教师模型,单步模型也能生成正确代码。

AI 中文摘要

语言生成几乎被普遍视为一个顺序过程:自回归模型一次输出一个token,而扩散语言模型则用一段长迭代优化轨迹替代token级别的串行性。本研究提出PlaidQ,一款用于代码生成的0.7B参数连续扩散语言模型,其轨迹可被大幅蒸馏至仅需少量去噪步骤——甚至一步,从而实现高效代码生成。PlaidQ将预训练自回归模型重新用作连续token嵌入上的双向去噪器。我们通过分布匹配对PlaidQ进行蒸馏以实现少步生成,通过配对轨迹监督实现一步生成。在模型规模匹配的情况下,PlaidQ在代码生成任务上可与离散扩散语言模型相媲美。蒸馏后改变了质量-计算权衡边界:16步的学生模型在HumanEval和MBPP+上的pass@10分别达到31.78和40.49,超过了采样512步的PlaidQ教师模型。在极端情况下,配对轨迹蒸馏在单步去噪下实现HumanEval上7.07的pass@1,生成功能正确的程序。综上,这些结果证明连续扩散是实现少步及一步代码生成的可行路径。广义而言,连续扩散不仅是语言的另一种表示形式,它还提供了一种接口,使语言模型能够继承连续扩散建模的加速与蒸馏机制。训练和推理代码及模型检查点可在该https URL获取。

英文摘要

Language generation is almost universally treated as a sequential process: autoregressive models emit one token at a time, while diffusion language models replace token-level seriality with a long trajectory of iterative refinement. In this work, we introduce PlaidQ, a 0.7B continuous diffusion language model for code generation, and show that its trajectory can be aggressively distilled into only a few denoising steps---or even one, enabling efficient code generation. PlaidQ repurposes a pretrained autoregressive model as a bidirectional denoiser over continuous token embeddings. We distill PlaidQ with distribution matching for few-step generation and paired-trajectory supervision for one-step generation. At matched model scale, PlaidQ is competitive with discrete diffusion language models on code generation. Distillation then shifts the quality--compute frontier: a 16-step student reaches 31.78 and 40.49 pass@10 on HumanEval and MBPP+, surpassing the same PlaidQ teacher sampled for 512 steps. At the extreme, paired-trajectory distillation achieves 7.07 pass@1 on HumanEval with a single denoising step, producing functionally correct programs. Together, these results establish continuous diffusion as a viable path to few-step and one-step code generation. Broadly, continuous diffusion is not merely another representation for language: it provides an interface through which language models can inherit the acceleration and distillation machinery of continuous diffusion modeling. Training and inference code and model checkpoints are available at https://github.com/pengzhangzhi/plaidq.

CommentsThis manuscript is withdrawn to address institutional disclosure requirements concerning the research resources used in this work

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑