发表机构
Oracle(甲骨文)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI数据中心冷却滞后问题,提出JCTI系统,利用调度器已知的作业类别信息预调配冷却液,在蒙特卡洛试验中减少56.4%热违规和60.2%超调,实现冷却与调度协同。
AI 中文摘要
GPU密集型AI数据中心必须依赖液体冷却运行,因为空气无法在这些功率密度下散发热量。然而,冷却回路本身对即将运行的工作负载一无所知;它们仅在传感器捕捉到温度上升后才加大流量,这一过程可能需要30至50秒。我们构建了作业类别热意图(JCTI)系统以缩小这一时间窗口。调度器已知晓作业即将到来及其所属类别;JCTI将该信息直接传递给冷却控制器,使其能够在热量出现前预先调配冷却液。我们从MLPerf GPU功耗轨迹中提取了每个作业类别的热特征,并根据阿里巴巴集群数据调整了到达模式。在超过120次配对蒙特卡洛试验中,与直接PI控制回路相比,热违规次数减少了56.4%,累计超调量降低了60.2%。随着AI数据中心向吉瓦级电网负荷发展,并伴随高度波动的功率摆动,热感知调度减少了突发需求,并改善了电网侧的负荷预测。冷却与调度作为两个独立系统已运行多年,尽管各自掌握对方所需的信息,JCTI将它们连接起来。
英文摘要
GPU-dense AI data centers need to run on liquid cooling as air simply cannot shed the heat at these power densities. Yet the cooling loops themselves are blind to what workloads are about to run; they crank up flow only after a sensor catches a temperature climb, which can take 30 to 50 seconds. We built Job-Class Thermal Intent (JCTI) to close that window. The scheduler already knows a job is coming and what class it belongs to; JCTI feeds that information straight to the cooling controller so it can stage coolant before the heat shows up. We pulled the thermal signatures for each job class out of MLPerf GPU power traces and tuned arrival patterns against Alibaba cluster data. Over 120 paired Monte Carlo trials the numbers come out to 56.4% fewer thermal violations and 60.2% less cumulative overshoot than a straight PI loop. As AI data centers evolving towards gigawatt grid loads with highly fluctuating power swings, thermally-aware scheduling reduces sudden demand and improves load prediction in grid side. Cooling and scheduling have been running as two separate systems for years despite each one knowing something the other needs, JCTI wires them together.
Comments6 pages, 5 figures, IEEE conference paper, first submitted to IEEE ICCSP 2026 02/02/2026