发表机构
Greenpixie Ltd.(绿精灵有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种估算云LLM推理每令牌能源、碳排放和水消耗的方法,区分输入输出令牌,结合测量与建模,为优化云和SaaS资源提供数据支持。
AI 中文摘要
我们描述了一种估算云托管大型语言模型(LLM)推理中每个令牌能源成本的方法,区分输入(预填充)和输出(解码)令牌。在广泛的基于文本的任务上,使用开放权重模型进行推理基准测试时,测量图形处理单元(GPU)的能源使用量。非GPU硬件的其余服务器能源贡献根据推理墙钟时间进行估算。使用贝叶斯线性回归来建模每个令牌的能源与LLM规模、请求流量和硬件部署配置之间的关系。对于规模和部署未知的专有前沿LLM,根据命名约定和性能先验将其分入大小桶,并使用蒙特卡洛方法对可能的LLM配置空间进行采样,以给出具有代表性的平均每个令牌能源和不确定性。我们还描述了如何利用这些能源测量来估算二氧化碳当量($\mathrm{CO_2\text{-}eq}$)排放(包括使用和隐含)、以及每次AI推理令牌消耗的水量。该方法提供了可操作的数据,有助于降低云和软件即服务(SaaS)中的成本、电力使用、$\mathrm{CO_2\text{-}eq}$排放和水的消耗。
英文摘要
We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a wide range of text-based tasks. The remaining server energy contribution from non-GPU hardware is estimated from the inference wall time. Bayesian linear regression is used to model the relationship between energy per token and LLM size, request traffic, and hardware deployment configuration. Proprietary frontier LLMs of unknown size and deployment are binned into size buckets based on naming conventions and performance priors, and the space of possible LLM configurations is sampled with Monte-Carlo methods to give a representative average energy per token and uncertainty. We also describe how these energy measurements can be used to estimate the carbon-dioxide equivalent ($\mathrm{CO_2\text{-}eq}$) emissions, both usage and embodied, and water consumed per token of AI inference. This methodology provides actionable data that enables reductions in cost, electricity usage, $\mathrm{CO_2\text{-}eq}$ emitted and water consumed in cloud and Software as a Service (SaaS).
Comments25 pages, 12 figures