AI 中文总结
该研究揭示LLM分块预填充调度可降低功率斜坡速率,随系统饱和度提升效益增大,经转化为电网调节备用采购问题后,可减少电网快速斜坡备用容量,为运营商提供无成本需求侧塑形工具。
AI 中文摘要
大型语言模型(LLM)推理服务是快速增长的电力负荷,从电网规划角度看,其功率动态仍未被明确表征。利用真实测量的GPU功率轨迹,我们表明,分块预填充调度(一种为降低延迟、已在生产级LLM服务中默认部署的技术)是可调节功率斜坡速率的可控旋钮,且无需改变峰值功率。与将长提示计算拆分为更小步骤应能平抑功率峰值的直观假设相反,峰值功率保持相对稳定,而平均斜坡速率却大幅下降。关键的是,这种斜坡速率优势并非该策略的固定属性:它随系统饱和度单调增长,我们在两个独立维度上验证了这一点:并发度(轻负载时平均斜坡降低7.0%,重负载时达34.6%)和长提示(“鲸鱼”)请求负载(低鲸鱼发生率时统计上无显著变化,高鲸鱼比例/规模时达42.6%)。我们将单GPU机制转化为电网运营量——调节备用采购,作为机会约束问题,使用直接重采样真实测量功率轨迹的无模型自举法提出并求解。在代表性运行点,这对应电网运营商需配置的快速斜坡备用容量减少20.3%至22.7%,覆盖95%至99.9%的可靠性水平。综上,这些结果为电网运营商提供了一种当前即可使用的无成本需求侧塑形工具,其效益在数据中心运行最热、电网压力最突出时最为显著。
英文摘要
Large language model (LLM) inference serving is a fast-growing electricity load whose power dynamics remain uncharacterized from a grid-planning perspective. Using real, measured GPU power traces, we show that chunked prefill scheduling, a latency-motivated technique already deployed by default in production LLM serving, is a controllable knob that regulates power ramp rate without touching peak power. Contrary to the intuitive hypothesis that splitting a long prompt's computation into smaller steps should flatten its power spike, peak power stays relatively the same while mean ramp rate falls substantially. Critically, this ramp-rate benefit is not a fixed property of the policy: it grows monotonically with system saturation, and we confirm this along two independent axes: concurrency (7.0% at light load to 34.6% at heavy load, mean-ramp reduction) and long-prompt ("whale") request load (from statistically flat at low whale incidence to 42.6% at high whale fraction/size). We translate this single-GPU mechanism into an operational grid quantity, regulation-reserve procurement, posed and solved as a chance-constrained problem using a model-free bootstrap directly resampling real measured power traces. At a representative operating point, this translates to an estimated 20.3-22.7% reduction in the fast-ramping reserve capacity a grid operator would need to provision, across reliability levels from 95% to 99.9%. Together, these results give grid operators a no-cost demand-shaping tool available today, whose benefit is largest precisely when data centers run hottest and grid stress is most salient.
Comments15 pages, 9 figures