arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34253cs.LGcs.CL

DreamingGoose:从自回归Transformer到双向循环扩散语言模型的分阶段蒸馏

DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models

Julian Boesch, Andrew Wee, Alexander Stranzl

首次发表
浏览论文内容

中文总结 AI 辅助

提出分阶段蒸馏方法,将自回归Transformer转换为双向循环扩散模型,发现检索能力难以转移,通过动态间隔课程可部分恢复,并报告了扩展至8B及代码模型的负面结果。

中文摘要 AI 辅助

预训练的自回归Transformer代表了计算方面的大量沉没投资。现有的转换方法通过改变架构(从注意力到循环)或目标(从下一词预测到去噪)来复用该投资,但从未同时改变两者。我们将1.7B和8B规模的Qwen3教师模型转换为无注意力、双向、门控增量规则扩散学生模型,分三个阶段进行,以便每个能力都可以追溯到保留或失去该能力的阶段。语言建模仅部分转移且仅在分布内;上下文检索无法转移。在一个多查询召回探针上,教师模型得分为0.34-0.58,而两个转换后的学生模型得分均为0.000,且仅扩散预训练无法恢复检索。在最后阶段引入检索课程,逐渐延长键值表与寻址查询之间的间隔,仅能随机恢复检索:在固定时间表下,三个种子中只有一个学会检索。仅在运行中的准确率估计保持高于阈值时推进间隔,对这三个种子均有效,在真实文本上保持有效,并原样扩展到8B规模,其中三个种子中有两个成功。第三个种子未在其固定的16k步预算内学会:检索在依赖于种子的步骤(另外两个种子分别为6.5k和11k)突然开启,因此固定预算可能截断较晚的运行。一个边界在所有干预中均存在:每个学会检索的模型在从未出现在检索情节中的令牌上得分为0.000,且一个每批次重新采样键和值令牌的实验臂表明这是覆盖限制,而非对特定绑定的记忆。另外,我们将一个7B代码模型转换为3:1循环-注意力块扩散混合模型,训练85k步,并报告两个负面训练结果。

英文摘要

Pretrained autoregressive Transformers represent a large sunk investment in compute. Existing conversion methods reuse that investment by changing either the architecture (attention to recurrence) or the objective (next-token prediction to denoising), never both. We convert Qwen3 teachers at 1.7B and 8B into attention-free, bidirectional, gated-delta-rule diffusion students in three stages, so that each capability can be traced to the stage that kept or lost it. Language modeling transfers only partially and in-distribution; in-context retrieval does not transfer. On a multi-query recall probe where the teachers score 0.34-0.58, both converted students score 0.000, and diffusion pretraining alone does not restore retrieval. A retrieval curriculum in the final stage, which gradually lengthens the gap between a key-value table and the queries that address it, restores it only stochastically: on a fixed schedule, one seed in three learns to retrieve. Advancing the gap only while a running accuracy estimate stays above a threshold works for all three of those seeds, holds on real text, and carries unchanged to 8B, where two of three seeds succeed. The third had not learned within its fixed 16k-step budget: retrieval switches on abruptly at a seed-dependent step (6.5k and 11k in the other two), so a fixed budget can cut a late run off. One boundary survives every intervention: every model that learns retrieval scores 0.000 on tokens that never appeared in a retrieval episode, and an arm that resamples the key and value tokens every batch shows this is a coverage limit, not memorization of particular bindings. Separately, we convert a 7B code model into a 3:1 recurrent-attention block-diffusion hybrid over 85k steps and report two negative training results.

发表机构

  • Purdue University(普渡大学)
  • Obit Research
  • State University of New York at Stony Brook(纽约州立大学石溪分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑