AI 中文总结
本文提出DEFT框架,联合优化多GPU应用的任务放置与DVFS配置,在CUDASTF中实现原型,可降低GPU能耗与EDP,同时保持性能接近最优基准。
AI 中文摘要
能效已成为现代高性能计算系统的首要关注问题,因为它直接决定了固定功耗预算下可实现的吞吐量。尽管动态电压频率调节(DVFS)是降低GPU能耗的有效机制,但现有运行时系统将DVFS与任务放置、GPU间通信解耦,仅聚焦单GPU执行,或无法适配多GPU环境下的任务粒度和运行时竞争。因此,当前调度器无法捕捉任务放置、频率选择与GPU间数据移动间的紧密耦合关系,而这种耦合从根本上决定了多GPU系统的能效-性能权衡。本文提出DEFT,一种感知能效的调度框架,针对基于任务的多GPU应用,联合优化任务到设备的分配及每个GPU的DVFS配置。DEFT采用成本模型驱动策略,整合松弛感知、吞吐量感知及任务执行成本、GPU间数据移动、DVFS转换开销的显式建模,可在动态运行时条件下,以任务粒度实现放置与频率决策的协同。我们在CUDASTF运行时中实现DEFT原型,在五个优化目标上验证其有效性。评估结果显示,DEFT在NVIDIA L40S和L4上平均分别降低能耗14.8%和4.8%,降低EDP(能量延迟积)9.9%和3.7%,同时保持性能与最快基准的差距在1.5%以内。
英文摘要
Energy efficiency has become a first-order concern in modern high-performance computing systems, as it directly determines achievable throughput under fixed power budgets. Although Dynamic Voltage and Frequency Scaling (DVFS) provides an effective mechanism for reducing GPU energy consumption, existing runtime systems decouple DVFS from task placement and inter-GPU communication, focus on single-GPU execution, or cannot adapt frequency to task granularity and runtime contention in multi-GPU environments. Consequently, current schedulers fail to capture the tight coupling between task placement, frequency selection, and inter-GPU data movement that fundamentally governs energy-performance trade-offs on multi-GPU systems. This paper presents DEFT, an energy-aware scheduling framework that jointly optimizes task-to-device assignment and per-GPU DVFS configuration for task-based multi-GPU applications. DEFT employs a cost-model-driven strategy that integrates slack awareness, throughput awareness, and explicit modeling of task execution cost, inter-GPU data movement, and DVFS transition overheads, enabling coordinated placement and frequency decisions at task granularity under dynamic runtime conditions. We prototype DEFT within the CUDASTF runtime and demonstrate its effectiveness across five optimization objectives. The evaluation shows that DEFT reduces energy consumption by 14.8% and 4.8% on average on NVIDIA L40S and L4, and reduces EDP by 9.9% and 3.7%, respectively, while maintaining performance within 1.5% of the fastest baseline.
CommentsPublished in the Proceedings of the 40th ACM International Conference on Supercomputing (ICS 26)