发表机构
Vizuara(Vizuara)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种基于精确提示长度、预测输出长度、KV缓存压力和SLO类别的学习型请求路由策略,经校准后在高负载下实现最高有效吞吐量,并显著降低硬件资源需求。
AI 中文摘要
分离式LLM服务将计算密集的预填充阶段和内存密集的解码阶段分别部署在独立的GPU池上。DistServe、Splitwise和Mooncake等系统实现了这种分离的高效运行,但路由仍然决定了每个请求由哪个实例处理。我们研究了一种路由器,它利用精确的提示长度、预测的输出长度、准入后的KV缓存压力以及SLO类别来估计每个实例上的额外完成时间。我们在离散事件模拟器中开发了该策略,并在八块NVIDIA A40 GPU上进行了验证,每块GPU运行一个vLLM引擎,使用NIXL在池之间传输KV缓存。所有工作负载均在实测饱和状态下运行。在三条混合的突发到达轨迹中,校准后的路由器实现了最高的平均有效吞吐量,达到0.864,而轮询、最少负载和长度启发式方法的有效吞吐量在0.835到0.847之间。它还在各轨迹间表现出最低的方差。它在所有三条轨迹上均优于轮询和长度启发式方法,并在两条轨迹上优于最少负载方法。在第三条轨迹上,它落后0.003,处于运行间噪声范围内。硬件校准至关重要:从模拟器推导出的常数导致有效吞吐量损失4.5个百分点,并损失约40%的尾部延迟优势,使评分器几乎退化为仅进行队列计数。收益随解码池规模和流量异构性的增加而增长,但在仅包含三个实例的池中消失,此时队列计数通常已足够。在极端资源稀缺情况下,贪婪成本最小化会将请求集中在评分最低的实例上,而盲目分散的效果更好。使用校准成本后,学习型路由器仅用六块GPU即可达到与轮询使用七块GPU时相当的有效吞吐量。
英文摘要
Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools. Systems such as DistServe, Splitwise, and Mooncake make this separation fast, but routing still determines which instances handle each request. We study a router that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class. We develop the policy in a discrete event simulator and validate it on eight NVIDIA A40 GPUs, each running a vLLM engine, with NIXL transferring KV caches between pools. All workloads run at measured saturation. Across three mixed, bursty arrival traces, the calibrated router achieves the highest mean goodput at 0.864, compared with 0.835 to 0.847 for round robin, least loaded, and a length heuristic. It also shows the lowest variance across traces. It beats round robin and the length heuristic on all three traces and least loaded on two. On the third, it trails by 0.003, within run to run noise. Hardware calibration matters: simulator derived constants cost 4.5 goodput points and roughly 40 percent of the tail latency advantage, reducing the scorer to little more than queue counting. Benefits grow with decode pool size and traffic heterogeneity but disappear in pools with three instances, where queue counts are often enough. Under extreme scarcity, greedy cost minimization concentrates requests on the cheapest scored instance, and blind spreading performs better. With calibrated costs, the learned router matches the goodput of round robin using six GPUs instead of seven.