发表机构
Wuhan University; Shanghai Jiao Tong University; The Hong Kong University of Science and Technology; Damen Database Co., Ltd.; Central China Normal University(武汉大学; 上海交通大学; 香港科技大学; 达梦数据库有限公司; 华中师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散LLM服务中推理成本预测不准确的问题,提出去噪工作负载表面(DWS)模型,保留二维块-步骤结构,通过粗到细训练实现轻量级预测,显著降低预测误差并优化调度延迟。
AI 中文摘要
随着扩散大语言模型(dLLMs)能力的增强,它们正从研究环境走向实际服务场景,其中请求管理(如调度和资源分配)依赖于对每个请求推理成本的准确估计。然而,常见的成本代理指标在dLLMs中表现不足:输出长度忽略了单次前向传播可同时解除多个令牌的掩码,而去噪步数忽略了各步骤成本的异质性。我们观察到,块自回归生成机制在输出块和块内去噪步骤上诱导出二维执行结构,而这些代理指标将其压缩为标量,丢弃了表征成本所必需的信息。基于这一洞察,我们提出去噪工作负载表面(DWS),它保留这种二维块-步骤结构作为概率表面,以加权异质的每步成本。随后,我们设计了一种从粗到细的训练方案,使得轻量级仅提示预测器能够准确预测复杂的DWS。该预测器即使在单个CPU核心上也能高效运行,避免了与服务模型的GPU争用。由于DWS将请求相关的执行行为与部署特定的成本因素解耦,预测器无需重新训练即可跨硬件配置迁移。在实际服务实验中,DWS将成本预测误差相较于基于标量的预测器降低高达2.50倍,而DWS引导的最短作业优先调度器将在线聊天机器人的端到端延迟降低高达1.92倍。
英文摘要
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.