arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时钟扩散:高效半自回归连续扩散语言模型

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky, Ran Zilberstein, Marianne Arriola, Gilad Turok, Guanghan Wang, Volodymyr Kuleshov, Michael Elad

arXiv 2610.00894首次发表:更新:

发表机构

NVIDIA; Cornell University(英伟达; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出时钟扩散框架,通过位置相关噪声调度实现半自回归连续扩散语言模型,支持变长生成和键值缓存,在OpenWebText和GSM8K上达到最优性能,并引入高效采样器缓存抓取。

AI 中文摘要

近期关于离散数据的连续扩散工作已展现出与同类离散扩散模型相当的性能。然而,这些连续对应模型缺乏作为语言模型实际应用所必需的关键特性,即变长生成和对键值缓存的支持,并且它们仍落后于自回归和离散扩散质量的前沿水平。在本工作中,我们解决了这些局限性。为此,我们引入了一种模型参数化方法,利用位置相关的噪声调度来定义半自回归(SAR)连续扩散语言模型(DLMs)。结合高效的训练和采样算法,我们将此框架称为时钟扩散(Clock Diffusion),并提出了我们方法的两个特例:块生成和滑动窗口生成。随后,我们定义了ClockDLMs,这是一族基于滑动窗口时钟扩散的高斯DLMs,在OpenWebText上达到了最先进的扩散似然界,甚至超越了性能优异的块SAR离散扩散模型。在TinyGSM上训练的ClockDLMs也在GSM8K基准上大幅超越连续基线,并达到或超过可比的SAR离散扩散模型。最后,基于我们的参数化,我们提出了更高效的采样器,称为缓存抓取(Cache Grab),它借鉴了离散扩散中加速推理的技术,例如提交概率超过置信度阈值的令牌以及自推测解码,进一步提升了我们模型的质量和效率。

英文摘要

Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑