WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
WattGPU: 在未见过的GPU和LLM上预测推理功耗和延迟
机构 * Universidad Politécnica de Madrid(马德里理工大学) ; Zurich University of Applied Sciences(苏黎世应用科学大学)
AI总结 提出WattGPU模型,利用公开的LLM元数据和GPU规格,无需硬件访问即可预测平均GPU功耗和令牌间延迟,在未见过的GPU和LLM上误差低至3.4%,优于基线方法。
Comments Accepted at 1st Workshop on Sustainability and Resource-Efficiency of Artificial Intelligence @ IJCAI 2026