扩散语言模型中的并行性、关键窗口与分离性
Parallelism, critical windows, and separations among diffusion language models
- Harvard University(哈佛大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文首次从理论上证明,均匀和高斯扩散语言模型在并行采样效率上优于掩码扩散,其前向传播次数与分布的内在复杂度相关,而掩码扩散因关键窗口更窄而需要更多步骤。
AI中文摘要:
扩散大型语言模型(dLLMs)的一个广受欢迎的卖点是其并行能力:即能够比自回归模型(每个token需要一次前向传播)更高效地生成文本序列。然而,在dLLMs的众多竞争范式(从掩码扩散到均匀扩散再到高斯扩散)中,关于这些不同方案在并行性方面如何比较的原则性理解仍然有限。在这项工作中,我们启动了对这三种领先方法并行能力的细粒度比较,并证明了以下结论:- 均匀扩散和高斯扩散可以在前向传播次数与底层分布的“双全相关”(dual total correlation)成比例的范围内进行采样,这是一种内在复杂性的度量,可能远小于上下文长度。此前,已知只有掩码扩散才能实现这一点。- 对于某一族随机经验测度,我们证明使用均匀或高斯扩散进行采样需要且仅需要$\widetilde{\Theta}(\sqrt{d})$次前向传播,然而存在近似的分数预言机(score oracles),使得掩码扩散需要$\widetilde{\Omega}(d)$次前向传播。这建立了三种主流dLLM范式之间在并行性上的首个可证明的分离。与流行的直觉(即掩码扩散由于必须提交token值而更难并行化)相反,后一种分离实际上源于这样一个事实:掩码扩散采样中的关键窗口(critical windows)渐近地比均匀扩散和高斯扩散采样中的关键窗口更窄。
英文摘要:
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following: - Uniform and Gaussian diffusion can sample in a number of forward passes which scales with the dual total correlation of the underlying distribution, a measure of intrinsic complexity which can be much smaller than the context length. Previously, it was only known how to achieve this using masked diffusion. - For a certain family of random empirical measures, we show that $\widetildeΘ(\sqrt{d})$ forward passes are necessary and sufficient to sample using uniform or Gaussian diffusion, yet there exist approximate score oracles for which $\widetildeΩ(d)$ forward passes are needed for masked diffusion. This establishes the first provable separation in parallelism between the three prevailing dLLM paradigms. Contrary to popular intuition that masked diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in masked diffusion sampling are asymptotically narrower than those in uniform and Gaussian diffusion sampling.