arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11226cs.AIcs.SYeess.SY

用强化学习降低AI数据中心能耗:从单GPU到集群的大语言模型训练实测功率控制

Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

Eliseo Curcio

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型训练的GPU功耗问题,提出PPO元控制器调整生成参数,在单GPU及集群环境下降低功率违规、提升输出与能效,验证了超额订阅的可行性。

中文摘要 AI 辅助

强化学习后训练主导了现代语言模型的开发,但其在GPU硬件上的功耗行为尚未被明确表征,而数据中心采用工作负载无关的机制(静态上限和反应式节流)来管理GPU功率,这些机制会不加区分地降低硬件速度。我们在1到4个A100 GPU上对7B、14B和72B规模的模型进行GRPO训练,每半秒记录一次功耗遥测数据(共38万多个样本),并训练了一个PPO元控制器,该控制器会根据实测功率调整工作负载自身的生成参数。与完整的500步7B轨迹相比,该控制器将功率限制违规降低了89.8%,同时令牌输出增加了18.1%,能效(每兆瓦时令牌数)提高了26.2%。将相同的控制器系列部署到72B模型上时,得到了重复的无效结果,经诊断为在模型分片下,分组大小执行器失去了控制权。对执行器控制权的扫描显示,应用相同参数作为生成并发时,仍保留了17-22%的功率控制权,从而分离出占用率与吞吐量的原则;基于该执行器重建的控制器,在三个副本上控制了72B的实时推出生成工作负载:与静态安全基线相比,输出增加了35.7%,预算违规为2.27±1.08%,比无控制操作的违规减少了87.2%,在受约束的控制器中具有最佳的平均吞吐量和每令牌能耗,且自适应阈值规则在三种操作条件之一中与之匹配。在实际测量窗口下,原始72B的瞬态从半秒分辨率下的23.6%降至30秒时的1.6%,5分钟时为零;由16个GPU组成的集群在30秒及更长时间内显示零违规,峰值需求为铭牌容量的50-56%。对于该集群组合,在运营商验证的前提下,铭牌容量的两倍左右的超额订阅似乎是可行的。我们量化了经济和碳影响,并指定了一个低成本的运营商试点项目。

英文摘要

Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware indiscriminately. We instrument GRPO training with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s (380,000+ samples), and train a PPO meta-controller that adapts the workload's own generation parameters to measured power. Against the full 500-step 7B trace, the controller cuts power-limit violations by 89.8% while increasing token output by 18.1% and energy efficiency by 26.2% (tokens per MWh). Deployed live at 72B, the same controller family yields replicated null results, diagnosed as the group-size actuator losing authority under model sharding. An actuator-authority sweep shows the same parameters applied as generation concurrency retain 17-22% power authority, isolating an occupancy-versus-volume principle; a controller rebuilt on that actuator controls a live 72B rollout-generation workload across three replications: 35.7% more output than a static safe baseline at 2.27 +/- 1.08% budget violations, 87.2% fewer violations than uncontrolled operation, and the best mean throughput and energy per token among constrained controllers, with an adaptive threshold rule matching it in one of three operating conditions. Under realistic measurement windows the original 72B transients fall from 23.6% at half-second resolution to 1.6% at 30 s and zero at 5 min; a composed 16-GPU fleet shows zero violations at 30 s and longer, with peak demand at 50-56% of nameplate. For this fleet mix, roughly twofold oversubscription of nameplate appears feasible, subject to operator validation. We quantify the economic and carbon consequences and specify a low-cost operator pilot.

↑