arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DScale:基于自适应验证的块扩散投机解码扩展

DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification

Rongjian Chen, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu

arXiv 2609.37532首次发表:更新:

发表机构

Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Victoria University of Wellington; University of Macau(中国科学院深圳先进技术研究院; 中国科学院大学; 惠灵顿维多利亚大学; 澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DScale通过路径感知瓦片、动态验证长度分配和固定地址工作区,在保持草稿模型完整性的同时,显著提升高并发下块扩散投机解码的吞吐量并降低延迟。

AI 中文摘要

大型语言模型应用的不断增长需要高效的推理。在高并发场景下,块扩散投机解码面临验证填充、候选被拒绝以及可变前缀与固定形状图之间不兼容的问题。统一截断会牺牲可接受的令牌。我们提出DScale,保留草稿模型架构、权重和完整草稿长度。一个独立的112K参数预测器既不需要置信度校准,也不需要硬件速度曲线准备。路径感知瓦片减少了填充。动态验证长度(DVL)分配将评分前缀打包到原生验证容量的一半中。固定地址工作区在验证和接受过程中传播变化的边界,同时重用捕获的图。在A100-40GB上,张量并行度为1,Qwen3-8B和Qwen3-4B覆盖四个数据集和并发度8-32,重用每个目标模型的冻结预测器。在这些配置下,几何平均吞吐量提升分别比DFlash高43.9%和48.8%,比DSpark高22.2%和37.7%,比Domino高24.4%和32.0%,且请求延迟更低。累积消融实验表明,依次添加这三种机制会持续提高几何平均吞吐量,而预算调整改善了接受令牌的保留。GPU分析显示,在GSM8K上,完整的解码步骤时间相对于DFlash减少了30.8-52.5%。

英文摘要

Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑