arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04198cs.AIcs.LG

ALoDLM:自适应循环扩散语言模型

ALoDLM: Adaptively Looped Diffusion Language Models

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao, Rajat Koner, Jiaye Wu, Linghan Xu, Xuanbai Chen, Xiang Xu, Zheng Zhang, Jakub Zablocki, Nishant Sankaran, Yifan Xing

首次发表
浏览论文内容

中文总结 AI 辅助

ALoDLM通过token自适应的潜在循环计算解决扩散语言模型的计算-难度不匹配问题,在1.7B和8B规模下于11个基准中超越现有DLM和AR基线,并保持快速并行解码。

中文摘要 AI 辅助

扩散语言模型(DLMs)通过并行预测多个token实现快速生成,但其实际应用仍受限于与同等规模自回归(AR)模型相比持续存在的质量差距。我们将这一差距归因于计算-难度不匹配:在部分观测序列中,一些未知token易于预测,而另一些则需要显著更多的计算量。然而,现有的DLMs在每个去噪步骤中对所有未知位置施加统一的计算深度。我们提出ALoDLM,用token自适应的潜在循环取代统一计算。在每个去噪步骤中,ALoDLM迭代地细化潜在表示,并根据token难度分配计算。准备好提交的token作为离散上下文反馈,而未解决的token则通过额外的循环传递保留并进一步细化其潜在状态。为了联合学习token预测和计算分配,我们将逐token的计算调度公式化为潜在变量,并推导出条件负证据下界(NELBO)。我们在1.7B和8B参数规模下训练ALoDLM。在十一个基准测试中,ALoDLM在平均基准分数上均优于所有评估的DLMs及相应的AR基线,且在两个规模下均如此。ALoDLM还保留了快速并行解码,在优化的推理引擎下,在评估的自回归和扩散模型中实现了强大的质量-效率权衡。

英文摘要

Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed sequence, some unknown tokens are easy to predict, while others require substantially more computation. Existing DLMs nevertheless apply uniform computational depth to all unknown positions at each denoising step. We introduce ALoDLM, which replaces uniform computation with token-adaptive latent recurrence. At each denoising step, ALoDLM iteratively refines latent representations and allocates computation according to token difficulty. Tokens ready to commit are fed back as discrete context, while unresolved tokens retain and further refine their latent states through additional recurrent passes. To learn token prediction and computation allocation jointly, we formulate token-wise computation schedules as latent variables and derive a conditional negative evidence lower bound (NELBO). We train ALoDLM at 1.7B and 8B parameter scales. Across eleven benchmarks, ALoDLM outperforms all evaluated DLMs and the corresponding AR baselines in average benchmark score at both scales. ALoDLM also retains fast parallel decoding, yielding a strong quality-efficiency trade-off among evaluated autoregressive and diffusion models under optimized inference engines.

发表机构

  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Amazon AGI(亚马逊通用人工智能)
  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

↑