AI 中文总结
研究针对块级扩散大语言模型块间解码串行问题,提出无需训练的\flowblock{}并行解码框架,基于门控波前解码和异构波前打包机制,相比串行模型提升了每秒令牌数、降低延迟并提高准确率,优于基于训练的基线。
AI 中文摘要
基于块的扩散大语言模型(dLLMs)在块级别顺序解码,能有效跨块重用键值缓存,但块间解码严格串行。先前工作试图通过训练后方法解锁块间并行性,但加速效果有限且常降低准确性。我们发现自校正dLLMs提供了无需训练的替代方案:令牌到令牌(T2T)编辑可修复使用稍旧上游上下文起草的令牌,因此下游块仅需信息性草稿而非最终的前序块。这将块的最终性从硬依赖转变为调度资源。我们提出了\textbf{\flowblock{}},这是一个基于两种机制的无需训练的并行解码框架。(i)\emph{门控波前解码}仅在满足就绪门时才允许块进入有界波前,通过T2T编辑联合优化活动块,并在保留精确冻结前缀键值缓存重用的窗口化块因果掩码下按顺序提交块。(ii)\emph{异构波前打包}为每个请求分配一个独立波前,同时将异步窗口打包成密集、形状稳定的批处理前向传播。在不同基准测试中,\flowblock{}相对于两个串行块级dLLMs,即LLaDA - 2.1和LLaDA - 2.0,每秒令牌数(TPS)提高了2.95倍和4.01倍,同时延迟分别降低了53.6%和77.1%。它还将平均准确率提高了1.3个百分点。与基于训练的块间并行基线D2F相比,\flowblock{}实现了更高的准确率和高达16倍的批处理服务吞吐量。
英文摘要
Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly serial. Prior work has attempted to unlock inter-block parallelism through post-training methods, but achieves only modest speedups and often degrades accuracy. We observe that self-correcting dLLMs offer a training-free alternative: token-to-token (T2T) editing can repair tokens drafted with a slightly stale upstream context, so a downstream block requires only an informative draft rather than a finalized predecessor. This turns block finality from a hard dependency into a scheduling resource. We propose \textbf{\flowblock{}}, a training-free parallel decoding framework built on two mechanisms. (i) \emph{Gated Wavefront Decoding} admits blocks into a bounded wavefront only when a readiness gate is satisfied, jointly refines active blocks via T2T editing, and commits blocks in order under a windowed block-causal mask that preserves exact frozen-prefix KV caches reuse. (ii) \emph{Heterogeneous Wavefront Packing} assigns each request an independent wavefront while packing asynchronous windows into dense, shape-stable batched forwards. Across different benchmarks, \flowblock{} improves tokens per second (TPS) over LLaDA-2.1 and LLaDA-2.0, two serial block-wise dLLMs, by up to 2.95$\times$ and 4.01$\times$, while reducing latency by up to 53.6\% and 77.1\%, respectively. It also improves average accuracy by 1.3 points. Compared with D2F, a training-based inter-block-parallel baseline, \flowblock{} achieves higher accuracy and up to 16$\times$ higher batched serving throughput.