arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文塔转换保留生成能力而冻结保留知识:MoE大语言模型的低预算AR到扩散转换

Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs

Wentao Lu, Tianyu Zhu, Jesse Clark

arXiv 2610.02657首次发表:更新:

发表机构

Celeris(Celeris)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文比较了MoE大语言模型低预算下两种AR到扩散转换方法,发现冻结上下文塔比原地转换保留更多父模型知识,并在HumanEval上提升11.6倍。

AI 中文摘要

将预训练的自回归(AR)模型转换为扩散语言模型(dLLM)可以在不重新预训练新模型的情况下实现并行生成。已发表的转换方法在训练数据量上相差约三个数量级,并且未在统一协议下进行比较。我们比较了同一30B混合专家(MoE)父模型的两个转换,固定了语料库、监督令牌预算、可训练参数集和评估框架,每种转换采用各自的训练方案。原地转换使用去噪和表示对齐损失更新父模型权重的子集;冻结塔转换则通过交叉注意力,以父模型的冻结因果副本作为条件。使用1B训练令牌,冻结塔模型在HumanEval pass@10上得分为71.60,而原地转换模型得分为6.19,提升了11.6倍。在同一预算下,冻结塔模型还保留了父模型GSM8K得分的95%和MMLUP-Pro得分的99%。密集父模型实验重现了HumanEval上的差异。在约500M令牌的两塔设计中,冻结上下文塔比训练它保留了更多的MMLUP-Pro性能,而两者在HumanEval上的观测得分相似。我们的理论分析表明,在硬注意力掩码和每轮一个位置的从左到右承诺下,两类转换都包含AR父模型的精确采样器。在共享损失下,冻结消除了通过上下文状态的梯度贡献。此外,评估协议对已发表的500B令牌转换的得分在各任务上产生双向显著影响,而其AR父模型的得分变化不超过三分,因此比较dLLM需要统一协议。这些结果表明,在测试的低预算范围内,冻结塔配置比原地转换保留了更多的父模型生成性能。

英文摘要

Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent's GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion's scores in both directions across tasks, while its AR parent's scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent's generation performance than in-place conversion.

Commentsv2: author order updated

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑