提升能量受限本地推理与训练性能的一个简单技巧
One Simple Trick for Improving the Performance of Energy-Limited Local Inference and Training
查看机构详情
- ISTA(奥地利科学技术研究所)
- TU Wien(维也纳工业大学)
- RedHat AI(红帽人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出将交替计算与内存操作的工作负载分块,以平滑功率和温度尖峰,防止降频,在DGX Spark上实现高达2%的性能与能耗改进,在多GPU服务器上效果为1-2%。
中文摘要 AI 辅助
能源供应和散热是现代GPU部署面临的两大主要挑战。虽然这些问题通常在新数据中心建设的背景下讨论,但同样的约束也适用于小型消费级设备,例如DGX Spark。在以计算密集型任务(如矩阵乘法)与内存受限操作(如范数或交叉熵)交替为特征的工作负载中,计算密集型部分可能会达到功率和/或热限制并开始降频。在这篇短文中,我们展示了将工作负载分块为以更高频率交替计算和内存操作的较小部分,可以平滑这些功率和温度尖峰,防止降频,从而显著缩短墙钟时间并降低总能耗。我们展示了在DGX Spark上可以利用这一效果的几种场景,性能提升和能耗降低高达2%,并证明同样的现象也发生在约束较少的系统(如多GPU服务器)上,尽管效果显著减弱,仅为1-2%。
英文摘要
Energy supply and heat dissipation are two of the main challenges with modern GPU deployments. While typically discussed in the context of new datacenter constructions, the same constraints also apply to small form-factor consumer devices, such as the DGX spark. In workloads characterized by alternating compute-intensive tasks such as matmuls with memory-bound operations such as norms or cross-entropy, the compute-intensive parts might hit power and/or thermal limits and start throttling. In this short paper, we show that chunking the workload into smaller parts that alternate compute and memory in higher frequencies, these power and temperature spikes can be smoothed out, preventing throttling and resulting in considerably faster wall-clock time and reduced total energy consumption. We present several scenarios in which this effect can be exploited on a DGX Spark with up to 2% performance and energy improvements, and demonstrate that the same phenomenon also happens on less constrained systems, such as a multi-GPU server, albeit at significantly reduced effect size of 1-2%.