发表机构
National Technical University of Athens; Harvard University; TWT GmbH Science & Innovation; NIKI Ltd Digital Engineering; University of Ioannina; National and Kapodistrian University of Athens; Massachusetts Institute of Technology(雅典国家技术大学; 哈佛大学; TWT有限责任公司科学与创新部; 尼基数字工程有限公司; 约阿尼纳大学; 雅典国立卡波迪斯特里亚大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种无需直接测量的GPU级LLM推理能耗分析估算方法,适用于模型比较、绿色编码分析及设计时评估。
AI 中文摘要
大语言模型(LLM)推理的运行能耗正成为已部署AI系统环境足迹中日益重要的组成部分。然而,直接测量推理能耗通常需要硬件遥测、功率仪器或特定基础设施的监控,这限制了其在比较研究、早期系统设计和可持续性报告中的适用性。本报告提出一种分析式结构化、经实证校准的GPU级方法,用于在NVIDIA H100级加速器上估算LLM推理能耗,无需直接运行时测量。所提估算器结合参数缩放的Transformer FLOP核算、校准的内存流量因子,以及FP16/BF16张量核心计算和高带宽内存传输的硬件特定能耗系数,明确将提示预填充与自回归解码分离,可估算输入Token、输出Token及完整推理请求的能耗。该方法进一步将总能耗分解为计算、参数访问、键值缓存写入和注意力读取组件,可分析其随模型规模、上下文长度和生成Token数的缩放行为。所得估算值并非用于替代物理功率测量,而是提供透明、可复现且假设明确的近似值,适用于模型比较、绿色编码分析及LLM推理工作负载的设计时评估。
英文摘要
The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure-specific monitoring, limiting its applicability in comparative studies, early-stage system design, and sustainability reporting. This report presents an analytically structured, empirically calibrated, GPU-level methodology for estimating LLM inference energy on NVIDIA H100-class accelerators without direct runtime measurement. The proposed estimator combines parameter-scaled transformer FLOP accounting, calibrated memory-traffic factors, and hardware-specific energy coefficients for FP16/BF16 tensor-core computation and high-bandwidth-memory movement. It explicitly separates prompt prefill from autoregressive decoding, enabling energy estimates for input tokens, output tokens, and complete inference requests. The methodology further decomposes total energy into compute, parameter-access, key-value-cache write, and attention-read components, allowing the scaling behavior with model size, context length, and generated-token count to be analyzed. The resulting estimates are not intended to replace physical power measurements; rather, they provide transparent, reproducible, and assumption-explicit approximations suitable for model comparison, green-coding analysis, and design-time evaluation of LLM inference workloads.
Comments20 pages, 3 figures, 6 tables. Accepted for oral presentation at the GREEN-AI Workshop, co-located with ECML-PKDD 2026