推理工程帕累托图谱:哪些优化主导成本、质量与延迟前沿?
The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?
- Vizuara
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
通过测量54种配置并校准模拟器,构建成本、质量与延迟的帕累托图谱,发现组合优化更常达前沿,且最佳配置依GPU和约束而异。
中文摘要 AI 辅助
大语言模型推理优化在不同模型、GPU、提示词和质量指标上报告加速效果,这使得它们难以比较或组合。我们构建了一个成本、质量和延迟的帕累托图谱,以识别不同部署约束下的最佳配置。由于穷举测试不切实际,我们测量了在vLLM 0.12上运行的Qwen2.5-7B-Instruct在L4、A100和H100 GPU上的54种配置,并利用这些锚点校准了一个模拟器。该模拟器在锚定批大小下重现了测量结果,跨活动漂移低于1.5%。一项独立的质量评估在200道GSM8K问题(每个提示词含五个示例)上测试了FP16、AWQ 4bit、FP8权重和FP8 KV缓存。稀疏注意力仅在模拟中评估。在校准网格上,36种配置中有18种达到帕累托前沿。组合方法比单一方法更常达到前沿,15种组合中有9种,而21种单一方法中有9种。质量测试改变了获胜者。AWQ 4bit在L4上将每令牌延迟降至基线的0.34倍,但严格GSM8K准确率损失了5.9%,在采样不确定性范围内勉强未达到95%的质量下限。灵活的答案提取与FP16准确率匹配,表明损失来自格式而非算术。FP8权重在三种GPU上均保留了基线准确率的99.4%,延迟为基线的0.61至0.65倍,并出现在四个领域获胜者中的三个中。朴素的FP8 KV缓存保持了正常吞吐量,但200个问题中无一回答正确,这表明仅靠速度是不够的。在两种提示词设计下,n-gram推测解码测得为基线的0.90至0.98倍,在此技术栈上未增加任何收益。最佳选择取决于约束和GPU:H100在严格延迟上获胜,而A100在吞吐量和低成本(每百万令牌0.106美元)上获胜。
英文摘要
LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since exhaustive testing is impractical, we measure 54 configurations of Qwen2.5-7B-Instruct running on vLLM 0.12 across L4, A100, and H100 GPUs and use these anchors to calibrate a simulator. It reproduces measurements at anchored batch sizes, with cross campaign drift below 1.5 percent. A separate quality evaluation tests FP16, AWQ 4bit, FP8 weights, and FP8 KV cache on 200 GSM8K questions with five examples per prompt. Sparse attention is evaluated only in simulation. On the calibrated grid, 18 of 36 configurations reach the Pareto frontier. Combined methods reach it more often than individual methods, with 9 of 15 combinations versus 9 of 21 single methods. Quality testing changes the winners. AWQ 4bit reduces per token latency to 0.34 times baseline on L4 but loses 5.9 percent of strict GSM8K accuracy, narrowly missing the 95 percent quality floor within sampling uncertainty. Flexible answer extraction matches FP16 accuracy, suggesting the loss comes from formatting rather than arithmetic. FP8 weights retain 99.4 percent of baseline accuracy at 0.61 to 0.65 times baseline latency across all three GPUs and appear in three of four regime winners. A naive FP8 KV cache maintains normal throughput but answers none of the 200 questions correctly, showing why speed alone is insufficient. Under two prompt designs, n gram speculative decoding measures at 0.90 to 0.98 times baseline and adds no benefit on this stack. The best choice depends on the constraint and GPU: H100 wins for tight latency, while A100 wins for throughput and low cost at 0.106 dollars per million tokens.