arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04010cs.LG

通过离散扩散解锁大语言模型(LLM)中的无损加速

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhentin… 展开作者

Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出 Uno 模型,通过扩散蒸馏和 Ψ-Spec 采样器,在不牺牲质量的前提下实现 LLM 生成的无损加速,8B Uno 模型在多任务基准上优于多款同类模型。

中文摘要 AI 辅助

大语言模型(LLM)的成功很大程度上依赖于下一个 token 预测(NTP),但其自回归(AR)结构需要缓慢的顺序 token 生成。为克服这一瓶颈,我们提出扩散增强型大语言模型(diffusion-augmented LLMs),这是一类定义了 AR 模型分布、同时利用扩散从该分布中并行抽取多个 token 的新型模型。我们将这些模型的参数解耦为两组:使用标准 NTP 目标训练的 AR 权重,以及用于同时生成多个 token 的轻量扩散权重。扩散权重通过简单的扩散蒸馏(Diffusion Distillation)阶段学习,该阶段对现有 LLM 训练流程的开销可忽略不计。我们还提出了 Ψ-Spec,这一系列采样器可在固定上下文长度下实现无损加速和推理时缩放。与投机解码不同,我们的方法不需要单独的草稿模型;与扩散大语言模型(d-LLMs)不同,它在不牺牲底层 AR 模型质量的前提下加速生成。生成的模型名为 Uno,可从头训练或通过增强现有开放权重 AR LLM 构建。Uno 在所有评估的批量大小下均实现了比领先的投机解码方法更高的吞吐量,并且在基础 AR 模型上实现了最高达 3 倍的加速,包括在设备支持的最大批量大小下。值得注意的是,我们的 8B Uno 模型在智能体工具使用、编码和长上下文推理的所有评估基准上,均优于领先的开放 d-LLM(26B DiffusionGemma)和专有模型 Mercury 2。我们在以下网址发布代码和检查点:this https URL

英文摘要

Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce $Ψ$-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to $3\times$ speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/

发表机构

  • Institue of Foundation Models(基础模型研究院)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Cornell Tech(康奈尔科技学院)
  • Harvard University(哈佛大学)
  • Cerebras Systems(赛布拉斯系统公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑