固定状态,长程可达:常量大小缓存为大规模块扩散带来的收益
Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale
浏览论文内容
中文总结 AI 辅助
本文提出利用状态空间缓存实现O(1)内存的块扩散解码,在256k上下文下相比注意力缓存降低11倍内存、提升14倍聚合吞吐,并支持8-16倍训练长度的检索。
中文摘要 AI 辅助
扩散语言模型并行解码词元,但其双向去噪器排除了快速自回归推理背后的朴素键值(KV)缓存。块扩散通过逐块解码恢复了缓存,且迄今为止部署在其上的块缓存与注意力机制绑定:内存为O(L),若作为免训练改造使用,则仅是模型计算的一种近似。这两个约束都可以被克服:将已完成块汇总为可复用状态的序列混合器支持块缓存,相应的块因果训练目标使缓存变得精确。我们在大规模下研究这一方案,在单一前沿目标下,用300B词元预训练三个3B块扩散去噪器(注意力、Mamba和混合架构),并通过单一缓存接口解码全部三个模型。只有状态空间缓存在序列长度上是O(1):其内存和每步延迟在任何上下文长度下保持恒定,而注意力缓存仍为O(L)。在256k词元时(此时注意力缓存已增长至82GB和29毫秒/步),Mamba缓存实现了4.3倍更低的延迟、11倍更少的内存和2.6倍更高的单流吞吐量;且由于该占用是恒定的,它也能随批次扩展,达到14倍的聚合吞吐量,而注意力在单流之外无法运行。相同的线性状态偏向使Mamba和混合骨干网络能够在其训练长度的8-16倍范围内持续检索,而注意力的检索在2倍处即崩溃,且未测量到质量损失。
英文摘要
Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the naive key--value (KV) cache behind fast autoregressive inference. Block diffusion restores caching by decoding block-by-block, and the block caches deployed on it so far are tied to attention: O(L)in memory and, if used as training-free retrofits, only an approximation of the model's computation. Both constraints can be overcome: sequence mixers that summarize finalized blocks into a reusable state support block caching, and the corresponding block-causal training objective makes the cache exact. We study this recipe at scale, pretraining three 3B block-diffusion denoisers (attention, mamba, and hybrid) on 300B tokens under one single-frontier objective and decoding all three through a single cached interface. Only the state-space cache is O(1) in sequence length: its memory and per-step latency stay constant at any context length, while an attention cache remains O(L). At 256k tokens (where attention has grown to 82GB and 29 ms/step), the Mamba cache delivers 4.3x lower latency, 11x less memory, and 2.6x higher single-stream throughput; and because that footprint is constant it scales with batch as well, reaching 14x the aggregate throughput, where attention cannot run beyond a single stream. The same linear-state bias lets the Mamba and hybrid backbones keep retrieving out to 8-16x their training length, whereas attention's retrieval collapses at 2x, at no measured quality cost.
发表机构
- Mila(米拉研究所)
- Concordia University(康考迪亚大学)
- ServiceNow Research(ServiceNow研究院)
机构由 AI 辅助整理,请以论文原文为准。