AI 中文总结
本研究测试MFU能否作为LLM训练的GPU功耗预测器,通过近3000次单设备训练基准测试,发现计算密集型 workload 下基于MFU的线性模型适用,特定参数拟合可大幅降低误差。
AI 中文摘要
高保真性能仿真器是设计和配置高效AI系统的关键,但当前工具缺乏预测功耗的能力。已有的GPU功耗模型依赖硬件利用率计数器,而这些计数器需在 workload 实际运行后才存在。本研究评估模型FLOPs利用率(Model FLOPs Utilization,MFU)——一种将实际吞吐量与硬件峰值能力关联的分析性、软件定义指标——是否可作为LLM的便携式软件定义GPU功耗预测器。我们在6种GPU上对近3000次单设备训练运行进行基准测试,覆盖不同模型系列、数值精度、批量大小和上下文窗口长度。结果发现,只要 workload 是计算密集型(如生产环境中的LLM训练),基于MFU的线性功耗模型适用于所有测试的GPU;按(GPU、数据类型、批量大小)而非按GPU拟合,可将单元内平均误差从约10%降至约1%,达到跨重复测量的噪声下限。
英文摘要
High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption. Established GPU power models rely on hardware utilization counters, which do not exist until the workload has actually run. This work evaluates whether Model FLOPs Utilization (MFU)-an analytical, software-defined metric relating achieved throughput to peak hardware capability-can serve as a portable, software-defined predictor of GPU power for LLMs. We benchmark almost 3000 single-device training runs across six GPUs, covering different model families, numerical precisions, batch sizes, and context-window lengths. We find that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training. Fitting per-(GPU, dtype, batch size) instead of per-GPU drops the within-cell mean error from around 10% to around 1%, matching the cross-repeat measurement-noise floor.
CommentsAccepted at MASCOTS 2026