焦耳点:AI 推理的能量最优运行点
The Joule Point: an Energy-Optimal Operating Point for AI Inference
- Holon Institute of Technology(霍隆理工学院)
- Afeka Academic College of Engineering(阿费卡工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出焦耳点,即 GPU 推理中能量最优的功率上限(约为额定功率的 43-46%),可降低能耗约 25-33%,并通过 ELF 数据集验证,将能耗转化为可调度的资源。
AI中文摘要:
服务于 AI 推理的数据中心力求最大化 GPU 利用率,默认情况下以满功率运行其显卡:这虽然最大化了吞吐量并降低了延迟,但能效低下,每个推理所消耗的能量高于在更优运行点完成相同工作所需的能量。这种低效是调度选择,而非硬件限制。其原因是物理性的:对于给定工作负载,GPU 的板级功率沿其功率-性能曲线超线性上升,再加上一个固定的功率底限(只要作业运行就需要支付),因此每个推理的能量消耗在运行点(GPU,功率上限)上呈 U 形。该最小值我们称之为焦耳点,是大型 GPU 额定功率的 43% 至 46% 的上限;将其设为上限可将每个推理的能量消耗降低约四分之一至三分之一,而资本和延迟成本适中:每个请求运行速度约慢 1.2 倍,保持总吞吐量需要同样倍数的更多显卡。在负载下,焦耳点几乎对每种显卡类型都是常数,因此每种显卡类型的一个静态上限即可捕获几乎所有的节省(平均损失低于 1%),将先前系统运行的在线每作业搜索转变为一次性表征。我们以 ELF 为基础,这是一个密集的功率上限数据集,扫描了四块 GPU 上的 20 个推理模型;在功率预算下于车队模拟中重放该数据集,将每个作业限制在满足其截止日期所需的最低功率,每个服务的作业能耗降低 18% 至 45%。这些结果将数据中心能源重新定义为一种可调度的资源,操作员可以在满足服务目标的同时,根据预算、电价或碳信号进行调整。
英文摘要:
Data centers serving AI inference strive to maximize GPU utilization, running their cards at full power by default: this maximizes throughput and holds latencies down, but it is energy-inefficient, spending more energy per inference than the same work needs at a better operating point. That inefficiency is a scheduling choice, not a hardware limit. The cause is physical: for a given workload, a GPU's board power rises superlinearly along its power-performance curve, on top of a fixed power floor that is paid for as long as the job runs, so energy per inference is U-shaped in the operating point (GPU, power cap). The minimum, which we name the Joule Point, is a cap at 43 to 46 per cent of a large GPU's rated power; capping to it cuts energy per inference by roughly a quarter to a third at a modest cost in capital and latency: each request runs about 1.2 times slower, and holding aggregate throughput takes that same factor more cards. Under load, the Joule Point is nearly a per-card-type constant, so a single static cap per card type captures nearly all the saving (mean penalty under one per cent), turning the online per-job search that prior systems run into a one-time characterization. We ground this in ELF, a dense power-cap dataset that sweeps 20 inference models across four GPUs; replaying it in a fleet simulation under a power budget, capping each job to the least power meeting its deadline spends 18 to 45 per cent less energy per served job. These results recast data-center energy as a schedulable resource an operator can adapt to a budget, an electricity price, or a carbon signal while meeting its service targets.