发表机构
Univ. Lille, CNRS, Inria; Univ. Bordeaux, CNRS, LaBRI; Univ. Lille, CNRS, Inria, IUF(里尔大学、国家科学研究中心、法国国家信息与自动化研究所; 波尔多大学、国家科学研究中心、LaBRI; 里尔大学、国家科学研究中心、法国国家信息与自动化研究所、国际大学联盟)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对vLLM的注意力内核类型、前缀缓存和分块预填充三个配置选项,在多种语言模型和推理任务中进行大规模评估,分析其对能耗、性能和准确性的影响,发现配置选项影响显著且因模型和工作负载而异,模型选择主导全局,配置调整提供局部改进。
AI 中文摘要
大语言模型正在重塑软件开发和维护方式,通常使用vLLM等推理引擎进行生产部署。此前研究集中在模型架构和硬件加速,而推理引擎配置对能耗、性能和输出质量的影响尚不清楚。本文对vLLM的三个配置选项(注意力内核类型、前缀缓存和分块预填充)进行大规模对照研究。在5个开放权重语言模型和5种不同推理任务中评估所有配置组合,共9000次运行和93600次测量。分析能耗、延迟和准确性,研究配置选项与任务间的主效应和交互效应。结果表明所选配置选项显著影响能量和性能,主要受注意力类型和前缀缓存驱动,分块预填充在默认配置和评估工作负载下影响有限。这些影响高度依赖模型和工作负载,无通用最优配置。还表明模型选择主导全局权衡,配置调整沿帕累托前沿提供局部改进,且推理选项会影响模型准确性。
英文摘要
Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference engines such as vLLM, which can efficiently serve pre-trained, highly configurable models. While prior work has focused on model architectures and hardware acceleration, the impact of inference engine configuration on energy consumption, performance, and output quality remains poorly understood. In this paper, we present a large-scale controlled study of three selected vLLM configuration options: attention kernel type, prefix caching, and chunked prefill. We evaluate all combinations of these configurations across 5 open-weight LLMs and 5 diverse inference tasks, totaling $9,000$ runs and $93,600$ measures. We analyze energy consumption, latency, and accuracy, and examine both main effects and interaction effects between configuration options and tasks. Our results show that the studied configuration options significantly impact energy and performance, mainly driven by attention type and prefix caching, while chunked prefill has a limited effect under the default vLLM serving configuration and evaluated workloads. These effects are highly model- and workload-dependent, and no configuration is universally optimal. We further show that model choice dominates global trade-offs, while configuration tuning provides local improvements along the Pareto frontier. Unexpectedly, inference options can also affect model accuracy.
CommentsSubmitted at a conference