arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11542cs.AI

表征面向功率弹性AI训练的任务功率弹性

Characterizing Job Power Elasticity for Power-Flexible AI Training

Philip Colangelo, Charles Dawson, Shayan Sengupta, Ayse Coskun, Varun Sivaram

首次发表
浏览论文内容

中文总结 AI 辅助

本文首次系统表征LLM训练的任务功率弹性,提出功率灵活性指数(PFI)量化功率降低的性能代价,实验表明基于PFI的功率分配在30%降功率下可恢复63%的性能差距,为功率感知AI基础设施奠定基础。

中文摘要 AI 辅助

大语言模型(LLM)训练是现代数据中心中电力需求增长最快的来源之一,而电力可用性是持续扩展AI基础设施的主要瓶颈。使这些工作负载的功耗变得灵活,可以释放额外的电力以支持AI增长,限制电价上涨,并提高现有电网基础设施的利用率。然而,要实现这种灵活性,我们必须首先了解当GPU功率降低时,训练工作负载的性能如何变化。本文首次对LLM训练中的任务功率弹性(即吞吐量对功率降低的敏感性)进行了系统性表征。为量化弹性,我们引入了功率灵活性指数(PFI),这是一个归一化指标,用于量化功率降低的性能代价,并为基于SLA的功率灵活性提供控制原语。我们收集了在H200上进行的131次LLM训练运行数据(外加24次H200验证运行和34次匹配的H100运行),包括稠密模型和混合专家模型、预训练和微调任务,以及最多32个GPU。我们发现LLM训练任务表现出显著但可变的功率弹性,并识别出可在运行时预测PFI的遥测信号。最后,我们证明了基于PFI的功率分配在功率约束下能最大化总吞吐量(以token/秒计)。在30%的功率降低下,基于PFI的功率分配每个任务恢复约1.5k token/秒,占等权分配与具有完美信息的理想分配之间性能差距的63%。我们的研究结果确立了功率弹性作为训练任务可测量属性的地位,并为功率感知、电网响应的AI基础设施奠定了基础。

英文摘要

Large language model (LLM) training is among the fastest-growing sources of electricity demand in modern data centers, and power availability is a primary bottleneck to continued AI infrastructure growth. Making the power consumption of these workloads flexible could unlock additional power for AI growth, limit increases in electricity prices, and improve the utilization of existing grid infrastructure. However, to realize this flexibility, we must first understand how the performance of training workloads changes when GPU power is reduced. This paper presents the first systematic characterization of \emph{job power elasticity} (the sensitivity of throughput to power reductions) in LLM training. To quantify elasticity, we introduce the \emph{Power Flexibility Index (PFI)}, a normalized metric that quantifies the performance cost of power reductions and provides a control primitive for SLA-aware power flexibility. We collect data from 131 LLM training runs on H200 (plus 24 H200 validation runs and 34 matched H100 runs), including both dense and mixture-of-experts models, pretraining and fine-tuning tasks, and up to 32 GPUs. We find that LLM training jobs exhibit substantial but variable power elasticity, and we identify telemetry signals that predict PFI at runtime. Finally, we demonstrate that PFI-aware power allocation maximizes total tokens/second throughput under power constraints. Under a 30\% power reduction, PFI-aware power allocation recovers ~1.5k tokens/s per job, 63\% of the performance gap between an equal-weight allocation and an oracle with perfect information. Our results establish power elasticity as a measurable property of training jobs and provide a foundation for power-aware, grid-responsive AI infrastructure.

发表机构

  • Emerald AI

机构由 AI 辅助整理,请以论文原文为准。

↑