arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边缘AI的电池代价:移动设备上LLM推理的环境影响研究

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

Édouard Guégain, Tristan Coignion

arXiv 2609.11940首次发表:更新:

发表机构

Univ. Bordeaux; CNRS; Bordeaux INP(波尔多大学; 法国国家科学研究中心; 波尔多国立理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了移动设备上LLM推理的能耗、性能与准确性,发现端侧推理能效低于服务器,且大部分环境影响来自设备制造,挑战了本地AI更可持续的假设。

AI 中文摘要

生成式人工智能的快速普及引发了隐私、延迟和性能方面的担忧,这些担忧推动了向“本地优先”AI的转变,即推理在用户设备上而非远程云服务器上执行。这种范式也给电池供电的智能手机带来了显著的计算负载,可能缩短电池寿命并提高移动设备的整体更换率。本文对端侧大语言模型(LLM)推理的能耗、性能和准确性进行了系统性研究。我们评估了来自不同模型家族、规模和量化级别的18个模型,在两款现代智能手机和一台服务器上,使用各自部署场景下的最先进技术进行测试。我们测量了每个生成token的能耗、token间延迟、模型准确性和电池循环消耗。我们的结果表明:(i)端侧推理平均比批处理服务器推理的能效低3倍;(ii)量化位宽与每token能耗之间的关系是非单调的,在测试的两款智能手机上均存在能耗甜点;(iii)18个模型配置中有8个位于准确性和能效的帕累托前沿上,使从业者能够构建电池感知的模型路由器;(iv)现实的建模假设不允许本地推理在每token环境影响上低于批处理服务器推理,其中88%至90%的影响归因于设备隐含碳而非电力消耗。这些发现挑战了本地AI比云推理更可持续的前提,并激励在电池供电的移动平台上部署边缘AI时,需要采用上下文感知和生命周期感知的模型选择。

英文摘要

The rapid diffusion of generative artificial intelligence raises privacy, latency, and performance concerns that motivate a shift toward "local-first" AI, where inferences are performed on the user's device instead of on remote cloud servers. This paradigm also places a significant computational load on battery-powered smartphones, potentially shortening battery life and increasing the overall replacement rate of mobile devices. This paper presents a systematic study of the energy consumption, performance, and accuracy of on-device large language model (LLM) inference. We evaluate 18 models from different model families, sizes, and quantization levels, on two modern smartphones and on a server, using the respective state-of-the-art for such deployments. We measure the energy per generated token, inter-token latency, model accuracy, and battery-cycle consumption. Our results show that (i) on-device inference is on average 3 times less energy-efficient than batched server inference; (ii) the relationship between quantization bit-width and energy per token is non-monotonic, with energy sweet spots on both tested smartphones; (iii) eight out of 18 model configurations lie on the Pareto front of accuracy and energy-efficiency, allowing practitioners to build battery-aware model routers; and (iv) realistic modeling assumptions do not allow local inference to be less environmentally impacting per token than batched server inference, with 88--90% of that impact attributable to device embodied carbon rather than electricity consumption. These findings challenge the premise that local AI is more sustainable than cloud inference, and motivate the need for context-aware and life-cycle-aware model selection when deploying edge AI on battery-powered mobile platforms.

Comments12 pages. Currently under submission at a conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑