Alpha扩散语言模型:因子分解本身并非问题
Alpha Diffusion Language Models: Factorization Alone Is Not the Problem
- Applied AI Institute(应用人工智能研究所)
- AXXX
- Yandex Research
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对离散扩散语言模型并行生成中预测不一致的问题,提出AlphaDLM,通过序列级alpha损失训练,在TinyGSM和SDAR-1.7B上提升了精度-计算权衡。
AI中文摘要:
离散扩散语言模型可以并行生成多个词元,但减少去噪步骤可能导致预测不一致。标准交叉熵训练拟合条件词元边缘分布,而并行生成需要一致的联合预测。我们引入了Alpha扩散语言模型(AlphaDLM),通过序列级alpha损失进行训练,该损失在alpha趋近于零时恢复交叉熵,并在alpha等于一时具有联合模态最优。我们的分析刻画了目标函数和因子分解如何共同决定拟合分布。我们确定了在中间alpha值下保留多个有效补全同时排除无效词元组合的条件。在TinyGSM上训练后,我们的方法在GSM8K上仅用四次模型评估就达到了34.6%的准确率。我们进一步将方法扩展到SDAR-1.7B,并在代码和数学基准上进行了评估。这些结果表明,改变训练目标可以改善因子分解扩散语言模型的精度-计算权衡。
英文摘要:
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing alpha and has a joint-mode optimum at alpha one. Our analysis characterizes how the objective and factorization jointly determine the fitted distribution. We identify conditions under which intermediate alpha preserves multiple valid completions while excluding invalid token combinations. Trained on TinyGSM, our method achieves 34.6% accuracy on GSM8K with only four model evaluations. We further scale the method to SDAR-1.7B and evaluate it on code and mathematics benchmarks. These results show that changing the training objective can improve the accuracy-computation trade-off of factorized diffusion language models.