arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

部署掩码扩散大语言模型:基于真实硬件的特性分析与设计原则

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Farhana Amin, Sabiha Afroz, Mona Moghadampanah, Dimitrios S. Nikolopoulos

arXiv 2608.23807首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过真实硬件实验分析LLaDA-8B-Instruct等dLLM的部署特性,发现其请求步数离散、批处理可分摊CPU开销等规律,提出适配dLLM的部署设计原则。

AI 中文摘要

掩码扩散语言模型(dLLM)原则上能够比自回归(AR)模型更快地生成文本,因为它们可以同时对多个token进行去噪。最近已有系统开始构建dLLM的部署基础设施,但此前尚无研究先测量这些模型在真实并发部署负载下的表现。若在缺乏此类依据的情况下构建部署系统,可能会沿用AR部署的假设,而这些假设对dLLM可能并不适用。为填补这一空白,我们开展了dLLM部署特性分析,使用LLaDA-8B-Instruct模型搭配D2F(离散扩散强制)LoRA适配器,在单块NVIDIA H200 GPU上进行实验,评估数据集为GSM8K和HumanEval。我们报告三项核心发现:第一,请求难度(即请求所需的去噪步骤数)是离散而非连续的:请求分为11个固定的步数等级(178至29000),且我们测试的所有信号均无法在生成开始前预测该等级(最佳R²值为0.150);第二,生成预算低于320 token的短生成基准测试会低估部署方差,因为请求在延迟差异显现前就被截断;第三,单请求的挂钟时间中仅有24%用于GPU计算,其余均为CPU端调度开销。批处理的主要作用是分摊此开销:在批大小为16时,每去噪步骤共享一次前向传播,相比按请求调度的基线,吞吐量提升16.0倍。我们还从结构上论证输出质量不应随批大小下降,并明确了支撑该结论的三项假设;在单请求规模下,我们测得GSM8K准确率为74%至76%。最后,我们推导了泊松到达场景下固定填充同步批处理的批超时规则。综合上述结果表明,部署扩散语言模型需要在每个去噪步骤层面实现并行性,这与AR部署中准入、逐出与已共享前向传播的交互方式存在差异。

英文摘要

Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (AR) models, since they denoise many tokens at once. Recent systems have begun building serving infrastructure for dLLMs, but none first measure how these models behave under real, concurrent serving load. Serving systems built without this grounding risk carrying over assumptions from AR serving that may not hold for dLLMs. We characterize dLLM serving to close this gap, using LLaDA-8B-Instruct with a D2F (Discrete Diffusion Forcing) LoRA adapter on a single NVIDIA H200 GPU, evaluated on GSM8K and HumanEval. We report three findings. First, request difficulty, the number of denoising steps a request needs, is discrete rather than continuous: requests fall into 11 fixed step-count levels (178 + 29k), and no signal we test predicts the level before generation starts (best R2 = 0.150). Second, benchmarks with short generation budgets below 320 tokens understate serving variance, since requests are cut off before the latency spread appears. Third, only 24% of single-request wall-clock time is GPU computation; the rest is CPU-side dispatch overhead. Batching mainly helps by amortizing this overhead: sharing one forward pass per denoising step improves throughput by 16.0x at batch size 16 over a per-request-dispatch baseline. We also argue structurally that output quality should not degrade with batch size, stating three assumptions this rests on; we measure 74 to 76% GSM8K accuracy at single-request scale. Finally, we derive a batch-timeout rule for fixed-fill synchronized batching under Poisson arrivals. Together, these results show that serving diffusion language models needs parallelism at the level of each denoising step, which differs from AR serving in how admission and eviction interact with an already shared forward pass.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑