发表机构
The Hong Kong Polytechnic University; Hong Kong University of Science and Technology; University of Macau; Wuhan University(香港理工大学; 香港科技大学; 澳门大学; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ETCInfer提出一种热感知冷却-计算联合调度器,通过设定点、频率和微批的联合优化,在保证热安全和SLO约束下,降低LLM推理作业能耗高达33.1%,并大幅减少GPU降频。
AI 中文摘要
AI数据中心中的大型语言模型(LLM)推理在GPU服务与设施冷却之间形成了一个耦合控制问题。提高环境温度设定点可以降低冷却能耗和碳排放,但也会缩小热余量,导致GPU降频,并引发服务级别目标(SLO)违规。在本文中,我们研究了LLM推理的冷却-计算联合控制:在满足热安全和延迟SLO约束的同时,最小化每个作业的GPU加冷却能耗。我们提出了ETCInfer,一种能效热感知调度器,它在作业前选择计算机房空调(CRAC)设定点,并在执行过程中调整每个GPU的频率和微批大小。ETCInfer通过从遥测数据中校准GPU发热、机箱散热、CRAC功率以及预填充/解码延迟关系,构建了紧凑的物理信息控制模型。这些模型估计隐藏热状态和到达降频的时间,使调度器能够在执行动作之前评估能耗、温度和延迟。我们将这个联合设定点-频率-微批控制问题表述为部分可观测马尔可夫决策过程,并设计了ETCAdapter,一种基于学习的控制器,在热安全和SLO约束下最小化每个作业的能耗。我们将ETCInfer实现为典型推理和集群管理栈之上的协调层。基于真实轨迹的模拟和验证实验评估表明,ETCInfer将总作业能耗降低了高达33.1%,热降频暴露降低了高达92.9%,并且即使在环境温度高达$48^{\circ}\mathrm{C}$的情况下,SLO违规率也保持在0.7%以下。
英文摘要
Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU-plus-cooling energy while satisfying thermal safety and latency SLO constraints. We present ETCInfer, an energy-efficient, thermal-aware scheduler that selects a pre-job Computer Room Air Conditioner (CRAC) setpoint and adapts per-GPU frequency and micro-batch size during execution. ETCInfer builds compact physics-informed control models by calibrating GPU heat generation, chassis heat dissipation, CRAC power, and prefill/decode latency relations from telemetry. These models estimate hidden thermal states and time-to-throttle, enabling the scheduler to evaluate energy, temperature, and latency before applying an action. We formulate this joint setpoint--frequency--micro-batch control problem as a partially observable Markov decision process and design ETCAdapter, a learning-based controller that minimizes per-job energy under thermal safety and SLO constraints. We implement ETCInfer as a coordination layer over typical inference and cluster management stacks. Evaluation across real-trace simulation and validation experiments shows that ETCInfer reduces total job energy by up to 33.1%, thermal throttle exposure by up to 92.9%, and keeps SLO violation rates below 0.7% even at ambient temperatures up to $48^{\circ}\mathrm{C}$.
Comments16 pages