arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

本地部署大语言模型(LLMs)的能效:消费级硬件上的初步GPU功耗定量基准测试

Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware

Philipp M. Zähl, Elja Dalipaj, Anika Hennig, Timon Bayer

arXiv 2608.00008首次发表:更新:

AI 中文总结

本文针对9款1B至7B参数的开源LLMs,在RTX 4060Ti 16GB GPU上开展能效基准测试,发现模型架构、量化策略等影响能效,gemma3:1b等模型能效最优,7B-Mistral能效最差,还需区分令牌生成模式。

AI 中文摘要

由于隐私顾虑和本地推理的需求,大语言模型(LLMs)的本地部署正受到关注。然而,消费级硬件上的能源成本仍未得到充分表征,因为大多数基准测试仅关注准确率。本文在单块消费级GPU(RTX 4060Ti 16GB)上对9个开源LLMs(参数规模从1B到7B)开展了可复现的硬件级能源基准测试。使用Ollama推理引擎,通过nvidia-smi以2Hz的频率在固定提示集上采样GPU功耗。我们评估了平均/峰值功耗、每条提示的总能源(J/提示)、每个输出令牌的能源(J/令牌)以及吞吐量(tok/s)。研究结果表明,除原始参数数量外,模型架构和量化策略等因素也会影响能效。具体而言,gemma3:1b和llama3.2:1b实现了最低的能源成本(0.56 J/令牌和0.65 J/令牌)和最高的吞吐量(>170 tok/s)。相比之下,7B-Mistral模型的每令牌能源消耗是最高效模型的4.4倍。值得注意的是,qwen3.5:2b因内部推理时间延长而表现出异常高的每提示能源,这凸显了在效率指标中区分令牌生成模式的必要性。

英文摘要

The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. This paper presents a reproducible, hardware-level energy benchmark of 18 open-source LLMs (0.5B to 7B parameters) executed on a single consumer GPU (RTX 4060ti 16GB). Using the Ollama inference engine, GPU power draw was sampled at 2hz via nvidia-smi across a fixed prompt set. We evaluate mean/peak power, total energy per prompt (J/prompt), energy per output token (J/tok), and throughput (tok/s). Our findings suggest that factors beyond raw parameter count, including model architecture and quantization strategy, drive energy efficiency. Specifically, qwen2.5:0.5b and tinyllama:1.1b achieve the lowest energy cost (0.2747 J/tok and 0.3234 J/tok) and the highest throughput (>325 tok/s). In contrast, the 7B-Mistral model consumes up to 8.6x more energy per token than the most efficient model. Notably, qwen3.5:0.8b(on) exhibits anomalously high per-prompt energy due to extended internal reasoning, highlighting the need to distinguish between token generation modes in efficiency metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑