去噪块而非词元:基于分支词元实现的压缩连续扩散高效生成
Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization
浏览论文内容
中文总结 AI 辅助
提出分支潜在扩散(BLD),将词元序列压缩为块潜在变量并并行分支解码,大幅提升长文本生成效率,相比基线吞吐量提升6倍以上。
中文摘要 AI 辅助
扩散语言模型(DLMs)通过迭代并行精炼生成文本,相比自回归(AR)解码具有更高吞吐量的潜力。然而,大多数DLMs仍为每个词元维持一个生成状态,因此每一步去噪都需处理与输出序列等长的状态序列,限制了并行生成带来的吞吐量提升。连续扩散模型提供了额外的自由度:单个连续状态可表示多个词元,使扩散能够在更短的潜在序列上运行。我们提出分支潜在扩散(BLD),利用这一灵活性将1024词元的序列压缩为仅64个块潜在变量,实现了16倍的压缩。BLD将潜在压缩与分支词元实现相结合,其中每个潜在变量由局部AR分支解码,所有分支并行运行。由于强压缩使联合潜在生成变得困难,BLD分组生成潜在变量,每组条件于先前生成的潜在变量。在同一GPU上的端到端评估中,BLD相比相似规模的ELF-L基线,生成FLOPs减少超过80倍,吞吐量提升超过6倍。与AR基线相比,BLD实现超过6倍的吞吐量提升和超过4倍的延迟降低。尽管进行了压缩,BLD仍保持有竞争力的局部流畅性和多样性,尽管长程连贯性仍具挑战。总体而言,BLD表明将扩散从词元级状态转移到压缩潜在序列可显著提升长序列生成的效率。
英文摘要
Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emph{Branching Latent Diffusion (BLD)}, which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction. BLD combines latent compression with \emph{branching token realization}, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than $80\times$ and increases throughput by more than $6\times$ relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than $6\times$ higher throughput and more than $4\times$ lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- Writer AI Research(Writer AI 研究院)
- William & Mary(威廉与玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。