arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BlockServe:用于高吞吐量扩散语言模型服务的块粒度连续批处理

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

Yuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, Philip S. Yu

arXiv 2607.08930首次发表:更新:

AI 中文总结

研究针对扩散大语言模型服务的收敛异质性问题,提出BlockServe框架,集成块粒度调度、混合状态执行及计算感知准入控制器,在五个基准测试中实现高吞吐量,确立块粒度调度为高吞吐量离线dLLM推理基础。

AI 中文摘要

扩散大语言模型(dLLMs)的高效服务受到收敛异质性的阻碍:批处理多个请求时,不同序列以不同速率收敛,导致较快请求在较慢请求后停滞,产生计算气泡和尾部延迟。我们提出了BlockServe,这是一个连续批处理框架,它将块粒度调度(在块边界立即淘汰已完成请求)与混合状态执行相结合,通过收集-散射索引将双缓存和并行解码扩展到异构批处理。此外,一个计算感知准入控制器通过令牌预算补充来扩展有效批处理容量。在五个基准测试中的Dream和LLaDA上,BlockServe在生成质量可比的情况下,吞吐量比Fast-dLLM高1.9至10.6倍,确立了块粒度调度作为高吞吐量离线dLLM推理的基础。

英文摘要

Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency. We present BlockServe, a continuous batching framework that integrates block-grained scheduling -- immediately evicting completed requests at block boundaries -- with mixed-state execution that extends dual cache and parallel decoding to heterogeneous batches via gather-scatter indexing. Furthermore, a compute-aware admission controller expands effective batch capacity through token-budgeted refill. On Dream and LLaDA across five benchmarks, BlockServe achieves 1.9--10.6$\times$ throughput over Fast-dLLM with comparable generation quality, establishing block-grained scheduling as a foundation for high-throughput offline dLLM inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑