发表机构
University of the Balearic Islands(巴利阿里群岛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过测量七种现代计算机视觉架构在ImageNet-10k上的训练能耗,发现GFLOPs与能耗相关系数为0.85,EfficientNet在性能和效率间最佳平衡,指导能源受限环境下的架构选择。
AI 中文摘要
深度学习的快速发展显著增加了与模型训练相关的能耗,使得能源效率成为日益相关的设计标准。本研究实证测量了在ImageNet-1k分类任务上训练七种现代计算机视觉架构的能量变化,包括MobileNetV3-Small、MobileNetV3-Large、EfficientNet-B0、EfficientNet-B1、ViT-B/32、ConvNeXt-Tiny和ViT-B/16,使用了一个同质的10,000张图像子集(ImageNet-10k)和统一的40轮基线配置,在哥伦比亚大学生物信息学与计算生物学中心(BIOS)的两块NVIDIA Tesla P100 GPU上执行。能量通过NVML直接记录,并与每个模型的计算复杂度进行对比。结果显示,浮点运算次数(GFLOPs)与以千瓦时(kWh)为单位的能耗之间的皮尔逊相关系数为0.85,表明计算复杂度是能耗的一个强但非完美的预测指标:具有相似GFLOPs的架构由于其主要操作的硬件效率差异,能耗差异可达3.1倍。EfficientNet变体在分类性能(验证集Top-5最高达97.15%)和能源效率(0.457-0.611 kWh)之间提供了最佳平衡,而Vision Transformers在所评估的配置下表现出最高的相对能耗和较低的分类性能。这些发现为能源受限计算环境中的架构选择提供了指导。
英文摘要
The rapid growth of deep learning has substantially increased the energy consumption associated with model training, making energy efficiency an increasingly relevant design criterion. This study empirically measures the energy variation of training seven modern computer vision architectures, MobileNetV3-Small, MobileNetV3-Large, EfficientNet-B0, EfficientNet-B1, ViT-B/32, ConvNeXt-Tiny, and ViT-B/16 for the ImageNet-1k classification task, using a homogeneous 10,000-image subset (ImageNet-10k) and a uniform 40-epoch baseline configuration executed on two NVIDIA Tesla P100 GPUs at the Bioinformatics and Computational Biology Center of Colombia (BIOS). Energy was recorded directly via NVML and contrasted with the computational complexity of each model. The results show a Pearson correlation of 0.85 between floating-point operations (GFLOPs) and energy consumption in kWh, indicating that computational complexity is a strong but imperfect predictor of energy expenditure: architectures with comparable GFLOPs exhibited consumption differing by up to 3.1x due to differences in the hardware efficiency of their dominant operations. The EfficientNet variants offered the best balance between classification performance (Val Top-5 up to 97.15%) and energy efficiency (0.457-0.611 kWh), while Vision Transformers exhibited the highest relative energy consumption and lower classification performance under the evaluated configuration. These findings guide architecture selection in energy-constrained computing environments.
CommentsInternational Conference on Next-Generation AI Technologies (ICNGAIT) Shanghai, China, 5 pages, 1 table