arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09919cs.LGmath.OCstat.ML

深度学习中海森矩阵的特征值:对称性及其破缺的起源

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

Yossi Arjevani

AI总结:

本文提出一个统一框架,将深度学习中海森矩阵谱的簇和离群值现象归因于从高度对称参考配置的偏离及其破缺,并通过三层ReLU网络等实例验证,解释了多层结构中的谱特征。

AI中文摘要:

深度学习训练模型中海森矩阵的谱表现出一种持续的模式:特征值组织成不同的簇,包括一个接近零的大块和少数孤立的离群值。本文表明,当原始设置被理解为偏离一个邻近的、否则隐藏的高度对称参考时,这些谱现象的天然解释就会出现。修改,包括架构、数据分布或参数度量的变化,会暴露一个邻近的参考配置,其海森矩阵展现出丰富的对称性——这些对称性并非由权重对称性所解释。在那里,对称性使得对谱的精确描述成为可能,迫使高维核和多重性大的特征值出现。回到原始配置会打破海森矩阵的对称性,从而产生观察到的簇和离群值的层次结构。该框架在一定的普遍性下发展,并对三层ReLU网络进行了详细分析,应用于卷积、图、Transformer模型以及NTK。进一步表明,相同的机制在逐层海森矩阵和高斯-牛顿矩阵中产生类似的谱结构。

英文摘要:

Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference. Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries. There, symmetry enables a precise description of the spectra, forcing high-dimensional kernels and eigenvalues of large multiplicity. Returning to the original configuration breaks the Hessian symmetry and thereby produces the observed hierarchy of clusters and outliers. The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK. The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.

↑