发表机构
The Chinese University of Hong Kong, Shenzhen; The University of Hong Kong; Huawei Technologies Co., Ltd.(香港中文大学(深圳); 香港大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM服务输出长度重尾分布的调度难题,OUTLETS利用推测解码草稿表示构建轻量长度预测器,降低MAE并将短请求P99延迟减少34.8%。
AI 中文摘要
大语言模型(LLM)服务中输出长度的重尾分布给资源配置和集群调度带来重大挑战。尽管输出长度预测可缓解这些问题,但现有方法存在关键缺陷:外部代理模型会增加大量延迟且保真度通常有限,而基于内部状态的方法虽高效却依赖对当前模型状态的浅层探测。我们发现推测解码(SD)与长度预测之间存在结构关联:高级框架(如EAGLE-3)中草稿解码器生成的潜在表示编码了可预测生成长度的信号。基于该见解,我们提出OUTLETS(Output-Length Prediction from Speculative Decoding Backbones,基于推测解码主干的输出长度预测),它将推测主干重新用作感知轨迹的长度预测器。当草稿表示已为推测解码计算时,OUTLETS仅添加一个轻量级回归头,且实现了比所评估方法更低的平均绝对误差(MAE)。在饱和的非聚合服务下,OUTLETS的预测使标准调度策略能够优先处理较短请求并在解码实例间更均匀地分配请求,将短请求的第99百分位(P99)延迟降低了34.8%。
英文摘要
The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-length prediction can mitigate these issues, existing approaches have key drawbacks: external proxy models add substantial latency and often have limited fidelity, whereas internal state-based methods are efficient but rely on shallow probes of current model states. We identify a structural connection between speculative decoding (SD) and length prediction: latent representations produced by the draft decoder in advanced frameworks (e.g., EAGLE-3) encode signals that are predictive of generation length. Building on this insight, we introduce OUTLETS (Output-Length Prediction from Speculative Decoding Backbones), which repurposes the speculative backbone as a trajectory-aware length predictor. When its draft representations are already computed for speculative decoding, OUTLETS adds only a lightweight regression head and achieves lower MAE than the evaluated methods. Under saturated disaggregated serving, OUTLETS predictions enable standard scheduling policies to prioritize shorter requests and distribute requests more evenly across decoding instances, reducing short-request P99 latency by 34.8%.
CommentsAccepted to EMNLP 2026