arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.13516cs.LG

神经网络训练中优化器选择对能效与性能的分析

An Analysis of Optimizer Choice on Energy Efficiency and Performance in Neural Network Training

  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

Tom Almog

更新

AI总结:

本文通过360次实验分析了八种优化器在三个基准数据集上的能效与性能,发现AdamW和NAdam一致高效,SGD在复杂数据集上性能更优但排放较高,为平衡性能与可持续性提供了见解。

AI中文摘要:

随着机器学习模型日益复杂且计算需求不断增加,理解训练决策对环境的影响对于可持续AI发展至关重要。本文提出了一项全面的实证研究,调查神经网络训练中优化器选择与能效之间的关系。我们在三个基准数据集(MNIST、CIFAR-10、CIFAR-100)上进行了360次对照实验,使用了八种流行的优化器(SGD、Adam、AdamW、RMSprop、Adagrad、Adadelta、Adamax、NAdam),每个实验设置15个随机种子。利用CodeCarbon在Apple M1 Pro硬件上进行精确的能耗追踪,我们测量了训练时长、峰值内存使用量、二氧化碳排放量和最终模型性能。我们的发现揭示了训练速度、准确性和环境影响之间存在显著的权衡,这些权衡因数据集和模型复杂性而异。我们确定AdamW和NAdam是一致高效的选择,而SGD在复杂数据集上表现出更优的性能,尽管其排放量较高。这些结果为寻求在机器学习工作流中平衡性能和可持续性的从业者提供了可操作的见解。

英文摘要:

As machine learning models grow increasingly complex and computationally demanding, understanding the environmental impact of training decisions becomes critical for sustainable AI development. This paper presents a comprehensive empirical study investigating the relationship between optimizer choice and energy efficiency in neural network training. We conducted 360 controlled experiments across three benchmark datasets (MNIST, CIFAR-10, CIFAR-100) using eight popular optimizers (SGD, Adam, AdamW, RMSprop, Adagrad, Adadelta, Adamax, NAdam) with 15 random seeds each. Using CodeCarbon for precise energy tracking on Apple M1 Pro hardware, we measured training duration, peak memory usage, carbon dioxide emissions, and final model performance. Our findings reveal substantial trade-offs between training speed, accuracy, and environmental impact that vary across datasets and model complexity. We identify AdamW and NAdam as consistently efficient choices, while SGD demonstrates superior performance on complex datasets despite higher emissions. These results provide actionable insights for practitioners seeking to balance performance and sustainability in machine learning workflows.

补充信息

↑