压缩前诊断:面向大语言模型服务轨迹的预测独立瓶颈见证优化
Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LLM Serving Traces
浏览论文内容
中文总结 AI 辅助
针对LLM服务轨迹重放成本高及现有方法的局限,提出BPW框架,通过三阶段流程构建紧凑可靠的重放套件,在三个数据集上优于16种策略,获两项指标提升
中文摘要 AI 辅助
生产环境中的大语言模型(LLM)服务会生成数百万个多样化请求,这使得跨服务配置的全轨迹重放成本日益高昂。现有的轨迹缩减方法主要保留工作负载分布或代表性请求,但揭示瓶颈的工作负载可能稀少且不具代表性。此外,一个组件的证据无法弥补另一组件缺失的证据,而使用预测的瓶颈作为目标真值会导致循环评估。这些限制使得有必要保留每个瓶颈组件的证据,而非仅依赖工作负载的代表性。我们提出了瓶颈保留见证(BPW),这是一个用于构建紧凑且诊断可靠的LLM服务重放套件的质量约束框架。BPW首先使用响应无关的工作负载特征和封闭源侧测量执行工作负载候选提名,该阶段识别可能暴露调度器、预填充、解码或KV缓存瓶颈的工作负载。接着,覆盖优先序列构建将多组件提议组织为可复用超边,并优先处理薄弱和未覆盖的维度。最后,瓶颈真值验证仅通过直接目标系统测量得出预测独立的标签,验证结果确定满足每个组件的直接双见证要求的最早前缀。在BurstGPT、ServeGen和Mooncake上的实验表明,BPW以紧凑的工作负载集达到了验证门限,且优于16种策略,在平均前缀宏F1值和WBRC-AUC上分别实现了2.3%和16.3%的相对提升。阶段解析和敏感性分析证实了其三个阶段的不同贡献和局部稳定性。我们的代码公开于此https URL
英文摘要
Production LLM serving generates millions of diverse requests, making full-trace replay across serving configurations increasingly expensive. Existing trace reduction methods mainly preserve workload distributions or representative requests, but bottleneck-revealing workloads may be rare and non-representative. Moreover, evidence for one component cannot compensate for missing evidence in another, while using predicted bottlenecks as target truth creates circular evaluation. These limitations make it necessary to preserve evidence for every bottleneck component rather than rely on workload representativeness alone. We propose Bottleneck-Preserving Witnessing (BPW), a quality-constrained framework for compact and diagnostically reliable LLM serving replay suites. BPW first performs Workload Candidate Nomination using response-blind workload features and closed source-side measurements. This stage identifies workloads that may expose scheduler, prefill, decode, or KV-cache bottlenecks. Coverage-Priority Sequence Construction then organizes multi-component proposals as reusable hyperedges and prioritizes weak and uncovered dimensions. Finally, Bottleneck Truth Verification derives prediction-independent labels solely from direct target-system measurements. The verified results determine the earliest prefix satisfying the direct two-witness requirement for every component. Experiments on BurstGPT, ServeGen, and Mooncake show that BPW reaches the verified gate with a compact workload set and outperforms 16 policies, achieving relative improvements of 2.3% and 16.3% in Mean prefix Macro-F1 and WBRC-AUC, respectively. Stage-resolved and sensitivity analyses confirm the distinct contributions and local stability of its three stages. Our code is publicly available at https://github.com/llmllmllm/BPW
发表机构
- Central South University(中南大学)
- University of Technology Sydney(悉尼科技大学)
- Chongqing Normal University(重庆师范大学)
- East China Normal University(华东师范大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。