批量大语言模型(LLM)服务的请求级能耗归因
Request-Level Energy Attribution for Batched LLM Serving
中文总结 AI 辅助
本研究提出JouleShare框架,通过计算Shapley能耗建立请求级基准,用JCalib校准模型提升批量LLM服务的请求级能耗归因精度,验证了token比例归因的偏差及JCalib的有效性。
中文摘要 AI 辅助
批量LLM服务可提升吞吐量,但使能耗核算变得复杂。GPU功率遥测是聚合数据,而可持续性报告、费用回收和工作负载分析往往需要请求级能耗计费。现有的推理能耗基准报告模型级、阶段级或token级能耗,近期的碳核算工作从概念上提出了Shapley公平性理念,但两者均未提供实测的请求级基准值,因此实践中使用的核算规则与公平分配的偏差程度仍不明确。我们提出JouleShare,这是一个包含两个组件的归因框架:一个离线工具通过在vLLM下重放请求子集,采用可复现协议,整合GPU功率遥测,为每个请求计算精确的Shapley能耗,从而建立该基准值;一个轻量级校准模型JCalib则学习从廉价的请求特征中预测Shapley份额,以便在服务时使用。在16次模型/工作负载运行中,token比例归因与精确Shapley的差异在静态批处理下平均为0.440归一化L1,在连续批处理下为0.458,该差距在三款数据中心GPU上均存在。JCalib将此误差降至静态批处理下的0.116,连续批处理下的0.177,甚至低于无法在线使用的独立测量基准,同时保留了精确的批量能耗效率。采样Shapley将实测基准扩展至更大的组规模,此时差距依然存在,且单次离线校准仍是最准确的可部署规则。结果表明,在批量执行下,token归因并非边际能耗的可靠替代,而实测的Shapley基准可校准低成本请求特征以实现更公平的归因。
英文摘要
Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis often require request-level energy charges. Existing inference-energy benchmarks report model-, phase-, or token-level energy, and recent carbon-accounting work motivates Shapley fairness conceptually. Neither provides measured request-level ground truth, so how far the accounting rules used in practice deviate from a fair allocation has remained unknown. We present JouleShare, an attribution framework with two components. An offline harness establishes this ground truth by replaying request subsets under vLLM with a reproducible protocol, integrating GPU power telemetry, and computing exact Shapley energy for each request. A lightweight calibration model, JCalib, then learns to predict Shapley shares from cheap request features for use at serving time. Across 16 model/workload runs, token-proportional attribution differs from exact Shapley by 0.440 normalized L1 on average under static batching and by 0.458 under continuous batching, a gap that reproduces across three data-center GPUs. JCalib reduces this error to 0.116 under static batching and 0.177 under continuous batching, below even a standalone-measurement baseline that is unavailable online, while preserving exact batch-energy efficiency. Sampled Shapley extends the measured reference to larger group sizes, where the gap persists and a single offline calibration remains the most accurate deployable rule. The results show that token attribution is not a reliable proxy for marginal energy under batched execution, and that measured Shapley ground truth can calibrate low-cost request features toward fairer attribution.