测量驱动的单GPU LLM服务中主机CPU共置干扰的诊断与缓解(在多GPU服务器上)
Measurement-Driven Diagnosis and Mitigation of Host-CPU Co-location Interference in Single-GPU LLM Serving on a Multi-GPU Server
查看机构详情
- School of Software Technology, Zhejiang University(浙江大学软件学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对单GPU LLM服务中主机CPU共置干扰问题,提出测量驱动的诊断程序CoTail,通过核心路径尾部指数和核心尾部抑制指标,筛选风险、分析尾部、选择OS级保护并验证SLO,显著提升吞吐量并降低延迟。
中文摘要 AI 辅助
在GPU服务器中,主机CPU在LLM推理期间通常未被充分利用。共置CPU工作负载可以提高资源利用率,但也可能严重损害服务质量。现有工作主要改进LLM服务引擎或研究CPU-GPU边界延迟,对于外部CPU工作负载如何影响服务路径以及操作员应如何选择保护策略提供的指导有限。本文研究单GPU LLM服务中的主机CPU共置干扰。我们表明,观察到的主要问题不是GPU内核变慢,而是CPU工作负载放大了GPU工作提交之前CPU侧服务阶段的长尾效应。为了捕捉这一效应,我们引入了核心路径尾部指数(CPTI)和核心尾部抑制(CTS)。基于这些指标,我们构建了CoTail,一种测量驱动的诊断程序,用于筛选工作负载风险、分析服务阶段尾部、选择操作系统级保护,并验证解码SLO合规性。在我们的主要设置中,未受保护的nginx共置使吞吐量降低78.8%,TTFT增加429.5%,TPOT增加362.4%。CoTail引导的保护使nginx吞吐量提升高达4.4倍,并将TPOT降低4.5倍。在常见基线部署SLO下,CoTail满足所有12个oracle可行的保留案例,而Always-rt为10/12,Macro-only为11/12。它还将RT使用量从28个案例减少到22个,并将平均共租户减速从56.65%降低到51.21%。
英文摘要
Host CPUs in GPU servers are often under-used during LLM inference. Co-locating CPU workloads can improve resource use, but it can also seriously hurt serving quality. Existing work mainly improves LLM serving engines or studies CPU-GPU boundary delays. It gives limited guidance on how external CPU workloads affect the serving path and how operators should choose protection policies. This paper studies host-CPU co-location interference in single-GPU LLM serving. We show that the main observed problem is not slower GPU kernels. Instead, CPU workloads amplify long tails in CPU-side serving stages before GPU work is submitted. To capture this effect, we introduce the Core Path Tail Index (CPTI) and Core Tail Suppression (CTS). Based on these metrics, we build CoTail, a measurement-driven diagnostic procedure that screens workload risk, profiles serving-stage tails, selects OS-level protections, and validates decode SLO compliance. In our primary setup, unprotected nginx co-location reduces throughput by 78.8%, increases TTFT by 429.5%, and increases TPOT by 362.4%. CoTail-guided protections improve nginx throughput by up to 4.4x and reduce TPOT by 4.5x. Under a common-baseline deployment SLO, CoTail satisfies all 12 oracle-feasible held-out cases, compared with 10/12 for Always-rt and 11/12 for Macro-only. It also reduces RT usage from 28 to 22 cases and lowers mean co-tenant slowdown from 56.65% to 51.21%.