arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19659cs.LGcs.DC

FleetSieve:面向SLO感知的大语言模型集群配置的决策关键型性能分析

FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

Huang Cheng, Scott Zhang, Aubert Li

首次发表
浏览论文内容

中文总结 AI 辅助

FleetSieve是一种面向SLO感知的LLM集群配置的决策关键型性能分析方法,通过联合建模容量与尾部延迟选择测量,可减少GPU资源消耗,避免违反SLO,提升集群决策效率。

中文摘要 AI 辅助

为大语言模型(LLM)服务集群选择张量并行度(TP)和副本数量颇具挑战,因为性能并非随TP单调变化,且可行选择会随负载改变。详尽的性能分析可解决这种不确定性,但会测量大量不影响最终资源分配的配置。我们提出FleetSieve,它根据测量对资源耦合、SLO感知的集群决策的预期影响来选择测量。FleetSieve联合建模容量和尾部延迟,比较保守与乐观分配,并在两者剩余决策差距低于指定容差时停止。针对310亿参数的开放权重模型,在固定H100测量网格上,FleetSieve使用22200 GPU秒达到了最优聚合决策,在固定对比中比均匀随机性能分析少6.9%。在200个随机揭示顺序中,它相对于随机性能分析的平均节省为5.4%(95%自助法置信区间:3.5-7.2%)。在固定对比中,Chat任务的节省为21.5%,而FleetSieve在Code任务中未使用最少的GPU秒数。联合容量和尾部建模还避免了选择一个配置,其46.4秒完成的p99延迟违反了30秒的SLO。在16-GPU分配中,错误的稀疏性能分析决策会导致每秒最多损失1.93个请求,以及最大最小满足率最多损失12.4个百分点。边界重复和BurstGPT测量支持观察到的负载依赖型尾部延迟机制。

英文摘要

Choosing tensor-parallel (TP) degrees and replica counts for an LLM serving fleet is difficult because performance is not monotonic in TP and the feasible choice can change with load. Exhaustive profiling resolves this uncertainty, but measures many configurations that do not affect the final resource allocation. We present FleetSieve, which selects measurements according to their expected effect on a resource-coupled, SLO-aware fleet decision. FleetSieve models capacity and tail latency jointly, compares conservative and optimistic allocations, and stops when their remaining decision gap is below a specified tolerance. On a fixed H100 measurement grid for a 31B-parameter open-weight model, FleetSieve reaches the oracle aggregate decision using 22,200 GPU-seconds, 6.9% less than uniform random profiling in the fixed comparison. Across 200 random reveal orders, its mean saving over random profiling is 5.4% (95% bootstrap CI: 3.5-7.2%). The fixed-comparison saving is 21.5% for Chat, while FleetSieve does not use the fewest GPU-seconds for Code. Joint capacity and tail modeling also avoids selecting a configuration whose 46.4-second completion p99 violates a 30-second SLO. In a 16-GPU allocation, an incorrect sparse-profile decision loses up to 1.93 requests/s and 12.4 percentage points of max-min fulfillment. Boundary repeats and BurstGPT measurements support the observed load-dependent tail-latency mechanism.

发表机构

  • Meta

机构由 AI 辅助整理,请以论文原文为准。

↑