arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2303.16464cs.LG

损失函数的Lipschitz性对采用Adam和AdamW优化器训练的深度神经网络泛化性能的影响

Lipschitzness Effect of a Loss Function on Generalization Performance of Deep Neural Networks Trained by Adam and AdamW Optimizers

  • aut.ac.ir(德黑兰大学(注:此处aut.ac.ir为伊朗德黑兰大学的域名缩写,根据上下文推断为作者所属机构))

机构由 AI 辅助整理,请以论文原文为准。

Mohammad Lashkari, Amin Gheibi

更新

AI总结:

该研究理论证明损失函数的Lipschitz常数是影响Adam/AdamW训练的深度神经网络泛化误差的重要因素,结合跨分布年龄估计实验验证了低Lipschitz常数、低最大值的损失函数可提升模型泛化能力,为损失函数选择提供指导。

AI中文摘要:

深度神经网络在优化算法层面的泛化性能是机器学习领域的核心关注问题之一,该性能会受到多种因素的影响。本文从理论上证明,损失函数的Lipschitz常数是降低采用Adam或AdamW训练得到的输出模型泛化误差的重要因素,该结果可作为优化算法为Adam或AdamW时选择损失函数的指导准则。此外,为在实际场景中评估该理论界,我们选取了计算机视觉中的人类年龄估计问题,为更好地评估泛化能力,训练集与测试集取自不同分布。实验评估表明,具有更低Lipschitz常数和最大值的损失函数能够提升采用Adam或AdamW训练的模型的泛化能力。

英文摘要:

The generalization performance of deep neural networks with regard to the optimization algorithm is one of the major concerns in machine learning. This performance can be affected by various factors. In this paper, we theoretically prove that the Lipschitz constant of a loss function is an important factor to diminish the generalization error of the output model obtained by Adam or AdamW. The results can be used as a guideline for choosing the loss function when the optimization algorithm is Adam or AdamW. In addition, to evaluate the theoretical bound in a practical setting, we choose the human age estimation problem in computer vision. For assessing the generalization better, the training and test datasets are drawn from different distributions. Our experimental evaluation shows that the loss function with a lower Lipschitz constant and maximum value improves the generalization of the model trained by Adam or AdamW.

补充信息

↑