发表机构
The Hong Kong Polytechnic University; Wuhan University; The Hong Kong University of Science and Technology(香港理工大学; 武汉大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HeatCache利用AIO液冷回路作为热缓冲,通过热预算调度LLM推理请求,在保证SLO和热安全下降低能耗,实现最高18%节能和81.7%热节流减少。
AI 中文摘要
LLM推理日益部署在机构级边缘以满足服务需求。然而,多GPU推理消耗大量电力并产生大量热量。为提高可持续性,运营商和法规通常要求提高环境设定点以减少冷却电力。这可能导致热节流和硬件老化加剧,从而引发服务级别目标(SLO)违规。本文提出HeatCache,一种热感知、节能的LLM推理调度器,适用于商业机箱级AIO液冷GPU在可持续环境温度下的运行。HeatCache将AIO回路视为临时热缓冲,通过热预算衡量,并基于来自HeatiTS的电感知热需求估计,调度请求以在热安全和SLO约束下最小化能耗。我们在vLLM之上实现HeatCache,并表明它最多可降低计算能耗18.0%,减少热节流暴露81.7%,即使在高达48°C的环境温度下,SLO违规率也保持在0.9%以下。
英文摘要
LLM inference is increasingly deployed at institution-scale edges to meet service requirements. However, multi-GPU inference consumes a large amount of electricity and produces substantial heat. To improve sustainability, operators and regulations often demand raising the ambient setpoint to reduce cooling electricity. This can increase thermal throttling and hardware aging, leading to Service-Level Objective violations. In this paper, we present HeatCache, a thermal-aware, energy-efficient LLM inference scheduler for commercial chassis-level AIO liquid-cooled GPUs at sustainable ambient temperatures. HeatCache treats AIO loops as a temporary heat buffer, measured by heat budget and schedules requests to minimize energy subject to thermal safety and SLO constraints, based on an electrical-informed heat-demand estimation from HeatiTS. We implement HeatCache atop vLLM and show that it reduces computing energy by up to 18.0%, decreases thermal-throttle exposure by 81.7%, and maintains SLO violation rates below 0.9% even up to $48~^{\circ}\mathrm{C}$.