AI 中文总结
本研究提出基于边际运行排放率(MOER)信号的碳感知大语言模型推理路由方法,通过多区域GPU测试平台实时验证,可减少50.9%的GPU归因运行排放,优于轮询及静态年平均MOER策略。
AI 中文摘要
大语言模型推理是一种快速增长的电力负荷,其边际碳强度在不同电网区域及一天内的变化幅度超过一个数量级,这使得请求放置成为一个极具吸引力的调控手段:无需重新训练,无需更改硬件。我们报告了基于边际运行排放率(MOER)信号在多区域GPU测试平台上进行碳感知推理路由的实时验证,该验证具有三项现有工作中不常见的特性:盲基准是实际基于生产压力的路由器,而非均匀放置;通过测量的并发曲线(来自NVIDIA DCGM的GPU遥测)归因的每请求能量,而非铭牌TDP;以及针对历史MOER而非驱动决策的预测对每个请求进行碳结算。核心实时结果是可行性:MOER信号在区域间引导推理,未观察到调度故障,作为生产路由器的严格且可逆的叠加。为评估效果规模,我们在电网多样化的美国大陆(CONUS)集群上重放了一年的每小时MOER。在主要建模配置中,与轮询放置相比,碳感知放置可将建模的GPU归因运行排放减少50.9%(95%块自举置信区间为48.5-53.3%)。由于重放是针对历史MOER而非预测进行调度,因此这是该配置下的上限;预测误差会降低实际运行的节约量。在会话固定之前,每小时最低MOER路由贡献约22.4个百分点,约为静态年平均MOER策略实现的54.0%放置减少量的40%。这些是针对一个集群和历史年份的建模结果,而非通用估计。我们还记录了一项实际观察:在比较区域时,应按绝对MOER而非百分位信号指数排名,该指数在每个区域内归一化,回答的是时间而非空间问题。
英文摘要
Large-language-model inference is a fast-growing electricity load whose marginal carbon intensity varies by more than an order of magnitude across grid regions and across the day, making request placement an attractive lever: no retraining, no hardware change. We report a live validation of carbon-aware inference routing on multi-region GPU testbeds driven by marginal operating emissions rate (MOER) signals, with three properties uncommon in prior work: a blind baseline that is an actual production pressure-based router rather than uniform placement; per-request energy attributed from GPU telemetry (NVIDIA DCGM) via measured concurrency curves rather than nameplate TDP; and carbon settlement of every request against historical MOER, not only the forecast that drove the decision. The central live result is feasibility: a MOER signal steered inference across regions with no observed dispatch failures, as a strict and reversible overlay on the production router. To size the effect, we replay a year of hourly MOER across a grid-diverse CONUS fleet. In the primary modeled configuration, carbon-aware placement reduces modeled GPU-attributable operational emissions by 50.9% versus round-robin (95% block-bootstrap CI 48.5-53.3%). Because the replay dispatches against historical MOER rather than a forecast, this is an upper bound under that configuration; forecast error would reduce operationally realized savings. Before session pinning, hourly lowest-MOER routing contributes about 22.4 percentage points, roughly 40% of the 54.0% placement reduction, beyond a static annual-mean-MOER policy. These are modeled results for one fleet and historical year, not a universal estimate. We also record a practical observation: when comparing regions, rank by absolute MOER rather than the percentile signal-index, which is normalized within each region and answers a temporal, not a spatial, question.
Comments18 pages, 4 figures. Live multi-region GPU validation plus a one-year historical marginal-emissions replay