arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

两层线性网络训练的全局指数收敛

Global Exponential Convergence of Two-Layer Linear Network Training

Stephen Y Zhang, Gabriel Peyré

arXiv 2610.09356首次发表:更新:

发表机构

Flatiron Institute; NYU; CNRS; ENS; PSL Univ.(熨斗研究所; 纽约大学; 法国国家科学研究中心; 巴黎高等师范学院; 巴黎文理研究大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明宽两层线性网络在丰富缩放下训练具有全局指数收敛,通过Bures流闭合动力学,并给出显式收敛速率及有限宽度稳定性,扩展至深度残差网络和动量情形。

AI 中文摘要

我们证明了在丰富缩放下,使用平滑的Polyak-Lojasiewicz预测损失训练的宽两层线性网络具有显式速率的全局指数(线性)收敛。因子中的梯度流恰好以神经元定律协方差的有限维Bures流的形式闭合,其中预测器动力学由隐藏协方差块进行预处理。当初始协方差满足谱支撑间隙条件时,均值场守恒定律为隐藏预处理块提供均匀谱下界。该条件包含正定性,同时仍允许奇异初始化。对于初始协方差$\Sigma_0 = \sigma^2 \mathrm{Id}$,损失以至少$4\sigma^2\kappa$的线性速率收敛到全局最小值,其中$\kappa$是PL常数。我们建立了该速率在有限宽度采样下的稳定性,以及因子梯度下降在依赖于平滑性、初始损失和守恒谱边界的显式步长区间内的全局收敛性。我们的论证逐层扩展到深度线性残差网络(ResNets),受残差路径界约束。在重球动量的情况下,训练动力学改为在位置和速度上以提升相位协方差的形式闭合。线性收敛在能量和阻尼的显式条件下成立,指定了允许阻尼的窗口。对于两尺度白初始化,当位置尺度足够大、初始损失差距和速度协方差固定时,该区间非空。数值实验说明了协方差几何,并比较了预测速率和观察速率。

英文摘要

We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses. Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predictor dynamics are preconditioned by hidden covariance blocks. Mean-field conservation laws provide uniform spectral lower bounds on the hidden preconditioning blocks when the initial covariance satisfies a spectral support gap condition. This condition encompasses positive definiteness while still allowing for singular initializations. For an initial covariance $Σ_0 = σ^2 \mathrm{Id}$, the loss converges to the global minimum with linear rate at least $4σ^2κ$, where $κ$ is the PL constant. We establish stability of this rate under finite-width sampling, as well as global convergence of factor gradient descent for an explicit stepsize interval depending on smoothness, the initial loss, and conserved spectral margins. Our argument extends layerwise to deep linear ResNets, subject to a residual-path bound. In the case of heavy-ball momentum, training dynamics close instead over positions and velocities in terms of a lifted phase covariance. Linear convergence holds under an explicit condition on the energy and damping, specifying a window of admissible dampings. For two-scale white initializations, this interval is nonempty for sufficiently large position scales, with a fixed initial loss gap and velocity covariance. Numerical experiments illustrate the covariance geometry and compare the predicted and observed rates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑