发表机构
Pimpri Chinchwad College of Engineering; Red Hat(平克里钦奇瓦德工程学院; 红帽公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GreenBench基准框架,在Apple M4 Pro上测试开源LLM的能效等,发现其能效远高于数据中心GPU,确定了不同场景下的最优模型。
AI 中文摘要
大语言模型(LLMs)的快速普及引发了人们对其推理阶段环境影响的担忧。尽管绿色人工智能研究已聚焦于数据中心GPU和嵌入式平台,但具备统一内存架构的Apple Silicon上的LLM推理的能耗情况仍未得到研究。本文提出GreenBench,这是一个基准测试框架,用于在配备48GB统一内存的Apple M4 Pro上,针对三个NLP任务评估五个开源LLM(参数规模为3-9B)的能效、吞吐量和碳足迹。该框架使用macOS powermetrics进行直接功耗测量,并利用Ollama的纳秒级精度计时。研究发现,M4 Pro在持续推理期间的CPU+GPU封装功耗仅为0.47W,系统总功耗为8-12W,在单用户部署中,其每token能效比数据中心GPU高30-40倍。较小模型(3-3.8B)的吞吐量比更大模型(7-9B)高2.6-4.2倍,每token能耗最多低62%。帕累托分析确定Qwen 2.5(7B)为准确率与能效的最优权衡,其MMLU得分为57%,吞吐量为59 tokens/s;而Llama 3.2(3B)适合对延迟要求严格的应用,吞吐量达175 tokens/s。本文还提供了封装级和系统级的每token能耗,以及针对印度和美国电网的二氧化碳排放估算。
英文摘要
The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM inference on Apple Silicon, with its unified memory architecture, remains unstudied. This paper presents GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs (3-9B parameters) across three NLP tasks on an Apple M4 Pro with 48 GB unified memory. Using macOS powermetrics for direct power measurement and Ollama's nanosecond-precision timing, we find that the M4 Pro draws only 0.47 W of CPU+GPU package power during sustained inference, with total system power of 8-12 W, achieving 30-40x better energy efficiency per token than datacenter GPUs in single-user deployment. Smaller models (3-3.8B) deliver 2.6-4.2x higher throughput and up to 62% less energy per token than larger models (7-9B). Pareto analysis identifies Qwen 2.5 (7B) as the optimal accuracy-efficiency trade-off at 57% MMLU and 59 tokens/s, while Llama 3.2 (3B) suits latency-critical applications at 175 tokens/s. We provide per-token energy at package and system levels with CO2 estimates for India and US grids.
Comments7 pages, 1 figure, 6 tables. Accepted at IEEE ICCUBEA 2026