可复现的大语言模型推理基准测试:用于回归测试的顺序隔离协议
Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing
浏览论文内容
中文总结 AI 辅助
提出顺序隔离协议,通过控制并发变化降低LLM推理基准测试方差,将平均CV从15.2%降至2.2%,并支持回归测试与成本建模。
中文摘要 AI 辅助
大语言模型(LLM)推理的可复现基准测试颇具挑战性,因为重复测量结果可能随执行状态和系统状态而变化。我们提出了顺序隔离方法论(Sequential Isolation Methodology),这是一种受控的基准测试与回归测试协议,旨在降低测量结果之间的运行间方差,同时有意改变工作负载并发度。我们在NVIDIA A100 80GB GPU上使用vLLM 0.9.1,对三个具有代表性的开源LLM进行了评估,覆盖六种上下文大小和八种并发级别,每种配置重复五次。最终协议将平均变异系数(CV)从受控程度最低的方法论阶段的15.2%降低至最终协议下的2.2%;基于每种配置五次重复的中位数(P50)首次令牌时间(TTFT)值计算的CV中,144个配置中的113个(78.5%)实现了低于3%的CV。测量结果还显示,在测试的堆栈上,200至500个并发用户之间出现了显著的延迟转变,并且三个模型的P99延迟存在描述性差异。我们还提供了一个显式的成本盈亏平衡模型,该模型对API定价具有敏感性。该协议旨在为可复现的比较和回归测试提供稳定的参考,而非预测不受控生产流量下的绝对行为。基础设施即代码(Infrastructure-as-Code)和基准测试脚本支持实验环境的复制。
英文摘要
Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.
发表机构
- Lucerne University of Applied Sciences and Arts(卢塞恩应用科学与艺术大学)
- Microsoft(微软)
- Uthereal AG
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。