arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06870cs.AR

G-Power:基于已知GPU聚合知识基础的架构级GPU功耗建模

G-Power: Architecture-level GPU Power Modeling with Aggregated Knowledge Foundations from Known GPUs

Qijun Zhang, Yao Lu, Shang Liu, Mengming Li, Chen Zhang, Dongbo Wang, Zhiyao Xie

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有架构级GPU功耗模型因采用微基准测试训练导致准确率低的问题,提出G-Power框架,通过利用已知GPU聚合知识的三阶段算法建模,在四款NVIDIA GPU上实现比AccelWattch更优的功耗预测性能。

中文摘要 AI 辅助

图形处理器(GPU)已成为大规模并行计算的关键计算资源。随着芯片复杂度不断提升,能效已成为现代GPU的重要设计目标。GPU功耗优化依赖快速功耗评估,需要架构级GPU功耗模型。然而,由于功耗标签采集耗时,现有模型仅采用简单微基准测试作为训练数据。微基准测试作为训练数据的局限性导致AccelWattch等现有架构级GPU功耗模型准确率较低。为解决微基准测试作为训练数据的局限性,本文提出G-Power,一种架构级GPU功耗建模框架,利用额外已知GPU芯片提供附加知识。G-Power利用额外已知GPU芯片的聚合知识基础,再对目标GPU进行微调。为提供额外已知GPU芯片的基础并捕捉相似性以利用这些基础进行微调,G-Power采用三阶段算法:1)用额外已知芯片进行预训练;2)类注意力聚合;3)对目标GPU进行微调。我们在四款现代NVIDIA GPU上评估G-Power,结果显示其准确率较高。G-Power平均可实现14%的低平均绝对百分比误差(MAPE)和0.88的高相关系数(R),相比AccelWattch,MAPE降低22%,R提高0.36。

英文摘要

Graphics Processing Units (GPUs) have been serving as critical computation resources for large-scale parallel computations. With increasing chip complexity, power efficiency has become an important design objective for modern GPUs. GPU power optimization relies on fast power evaluation, requiring architecture-level GPU power model. However, because of the time-consuming power label collection, only simple microbenchmarks are adopted for training. The limitation of microbenchmarks as training data incurs low accuracy for existing architecture-level GPU power models like AccelWattch. To address the limitation of microbenchmarks as training data, we propose G-Power, an architecture-level GPU power modeling framework that utilizes additional known GPU chips to provide additional knowledge. G-Power utilizes the aggregated knowledge foundation from additional known GPU chips and then performs fine-tuning on our target GPU. To provide foundations with additional known GPU chips and capture the similarity to utilize these foundations for fine-tuning, G-Power adopts a three-phase algorithm consisting of 1) pre-training with additional known chips, 2) attention-inspired aggregation, and 3) fine-tuning on our target GPU. We evaluate G-Power on four modern NVIDIA GPUs, demonstrating high accuracy. G-Power can achieve a low MAPE of 14% and a high correlation coefficient R of 0.88 on average, which are 22% lower MAPE and 0.36 higher R than AccelWattch.

补充信息

↑