AI 中文总结
该研究构建LLM推理分解能耗模型,在NVIDIA H100、H200 GPU上评估密集与MoE模型的能耗特征,发现需同时优化请求与令牌能耗,而非仅降低单位令牌能耗。
AI 中文摘要
大语言模型(LLM)推理服务按令牌计费,但GPU能耗是在推理窗口内消耗的,这种计费不匹配导致令牌归一化指标不完整,因为即使总请求能耗增加,平均输出令牌能耗也可能下降。我们通过分解能耗模型来表征这种行为:包含一次性的固定预填充成本和固定的生成设置成本,每个输出令牌生成步骤则增加边际步骤能耗。我们在NVIDIA H100和H200 GPU上针对密集模型和混合专家(MoE)模型评估该LLM推理能耗模型,报告请求能耗和令牌能耗作为模型类型(M)、阶段(P)、批大小(B)、上下文长度(C)和输出长度(N)的函数。对于H200上的Llama-3.2-1B,当批大小为16、上下文长度为4K时,输出长度从10令牌增加到512令牌,令牌能耗从7.46 J/令牌降至0.72 J/令牌,而总批处理推理窗口能耗从1.19 kJ增加到5.93 kJ。批处理也会降低令牌能耗,但增益受上下文限制:在输出令牌为10时,批大小16相比批大小1的增益,在上下文512时为6.31倍,在上下文4K时降至1.17倍。MoE模型放大了这种效应:稀疏路由和碎片化的专家执行在低并发时增加固定能耗,而批处理将该能耗分摊到更多生成的令牌上,大幅缩小了密集模型与MoE模型的令牌能耗差距。这些结果表明,感知能耗的服务应同时优化请求能耗和令牌能耗,而非仅降低每令牌能耗成本。
英文摘要
Large language model (LLM) inference serving is priced by tokens, but GPU energy is consumed over inference windows. This accounting mismatch makes token-normalized metrics incomplete, since average output-token energy can decrease even when total request energy increases. We characterize this behavior with a decomposed energy model: a fixed one-time prefill with a fixed generation setup cost, while each output-token generation step adds marginal step energy. We evaluate this LLM inference energy model on NVIDIA H100 and H200 GPUs across dense and mixture-of-experts (MoE) models, reporting both request energy and token energy as functions of model type (M), phase (P), batch size (B), context length (C), and output length (N). For Llama-3.2-1B on H200 at batch-16 and context-4K, increasing output length from 10 to 512 tokens reduces token energy from 7.46 to 0.72 J/token while total batched inference-window energy increases from 1.19 to 5.93 kJ. Batching also reduces token energy, but the gain is context-bounded: at 10 output tokens, the batch-16 to batch-1 gain falls from 6.31x at context-512 to 1.17x at context-4K. MoE models amplify this effect: sparse routing and fragmented expert execution increase fixed energy at low concurrency, while batching spreads that energy across more generated tokens and substantially narrows the dense-vs.-MoE token-energy gap. These results show that energy-aware serving should jointly optimize both request energy and token energy, rather than only reducing per-token energy cost.
CommentsAccepted at the 2026 IEEE International Symposium on Workload Characterization (IISWC 2026). 13 pages, 6 figures, 9 tables