宽度依赖模型超参数和$\ell_2$正则化对两层ReLU网络损失景观的影响
Effects of width-dependent model hyperparameters and $\ell_2$-regularization on the loss landscape of two-layer ReLU networks
浏览论文内容
中文总结 AI 辅助
研究两层ReLU网络中宽度依赖模型超参数和$\ell_2$正则化对损失景观的影响,推导超参数条件及解析解,通过实验发现AdamW可防参数坍缩,揭示了相关因素对损失景观几何结构的影响。
中文摘要 AI 辅助
理解深度神经网络仍是机器学习的核心挑战。特别是,即使是两层ReLU网络的理论特性,尤其是在存在权重衰减的情况下,仍知之甚少。为此,我们推导了超参数设置的充分条件,在该条件下全局最小值坍缩到零解。有趣的是,我们的实验表明,使用AdamW作为优化器可防止学习参数的坍缩,而使用SGD则不然,这可能有助于解释AdamW在深度学习训练中的成功。此外,当将输入维度限制为1时,我们推导了两层ReLU网络全局最优参数集的解析解,并表明$\ell_2$正则化对连通性具有宽度不变的影响,但其降维效果随着网络宽度的增加而变强。这些结果为宽度依赖的超参数如何影响正则化损失景观的几何结构提供了见解。
英文摘要
Understanding deep neural networks remains a central challenge in machine learning. In particular, the theoretical properties of even two-layer ReLU networks, especially in the presence of weight decay, remain poorly understood. To this end, we derive a sufficient condition on the hyperparameter settings under which the global minima collapse to the zero solution. Interestingly, our experiments reveal that using AdamW as an optimizer prevents the collapse of the learned parameters, whereas using SGD does not, which may help explain the success of AdamW in deep learning training. In addition, when restricting the input dimension to one, we derive an analytical solution for the globally optimal parameter sets of two-layer ReLU networks and show that $\ell_2$-regularization has a width-invariant effect on connectivity, but its dimensionality-reducing effect becomes stronger as the network width increases. These results provide insight into how width-dependent hyperparameters influence the geometry of regularized loss landscapes.
发表机构
- Okinawa Institute of Science and Technology(冲绳科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。