基于广义利普希茨光滑性的神经网络梯度下降收敛性保证
Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
浏览论文内容
中文总结 AI 辅助
该研究针对无特殊初始化或数据集要求的一般前馈神经网络,利用激活函数的广义利普希茨光滑性,证明了梯度下降的收敛性,给出了最小平方梯度范数的收敛速率。
中文摘要 AI 辅助
我们针对任意宽度或深度的一般前馈神经网络,在初始化或数据集无特殊要求的情况下,建立了梯度下降的收敛性保证。仅假设激活函数满足利普希茨光滑、利普希茨连续且线性有界的性质——该性质适用于线性、tanh、softplus和sigmoid激活函数。对于损失函数,要求其在模型输出上是利普希茨光滑的,这一点对于均方误差损失成立。关键理论洞见在于,激活函数的利普希茨性质即使经过多次复合也会部分保留,由此得到一种新颖的广义利普希茨光滑性条件:梯度的变化被参数空间的变化乘以两个端点处参数范数的多项式项所上界约束。该条件同时适用于模型函数和损失函数,由此可得到一个下降引理:只要学习率相对于参数范数足够小,损失就会下降。通过确保参数范数不会过快增长至无穷,我们证明,对于L层神经网络,最小平方梯度范数在T次迭代中以O(1/T^(1/L))的速率收敛至零。
英文摘要
We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Lipschitz smooth---properties that hold for linear, tanh, softplus, sigmoid, and smoothed ReLU functions---and that the loss function is Lipschitz smooth in the model outputs, a mild condition satisfied by mean squared error and binary cross-entropy loss. By relating the Lipschitz properties of one layer to the next, we obtain a novel generalized Lipschitz smoothness condition for an $L$-layer neural network where the change in gradient is upper bounded by the change in the parameter space, multiplied by a polynomial of the parameter norms at both endpoints of degree $2L - 2$. This yields a descent lemma where the loss decreases as long as the learning rate is small enough with respect to the parameter norms. By ensuring that the leading term of the polynomial grows at a controlled rate, we prove that the minimum squared gradient norm converges to zero in $T$ iterations at rate $O(1/T^{1/L})$.