arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经网络的海森谱如何依赖于数据

How the Hessian-Spectrum of Neural Networks Depends on Data

Jasraj Singh, Enea Monzio Compagnoni, Antonio Orvieto

arXiv 2607.13631首次发表:更新:

AI 中文总结

研究神经网络海森谱与数据的关系,推导线性网络海森矩阵特征值,发现分类任务中解的锐度与样本类别比例有关,经实验验证预测并分析相关影响,结论适用于更实际学习设置。

AI 中文摘要

海森矩阵在研究深度学习中的损失景观、优化动力学以及设计泛化度量、二阶学习算法等方面是一个重要的研究对象。以往的工作主要集中在实证结果或在过于简化的设置下进行理论处理。在这项工作中,我们推导了具有任意宽度和深度的线性网络以及具有任意数量样本、特征和标签的数据集的海森矩阵的特征值。重要的是,对于具有均方误差损失的分类任务,我们发现解的锐度与属于任何类别的样本的最大比例直接相关。我们通过实验验证了我们的预测,并系统地分析了一次去掉不切实际假设以及引入非线性的影响。我们观察到我们的预测在大多数情况下相当稳健,这使我们能够将结论扩展到更实际的学习设置中。

英文摘要

The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of generalization, second-order learning algorithms, etc. Prior works have focused on empirical results or pursued a theoretical treatment under overly simplified settings. In this work, we derive the eigenvalues of the Hessian of linear networks with arbitrary widths and depths, and datasets with an arbitrary number of samples, features, and labels. Importantly, for classification tasks with MSE loss, we identify that the sharpness of the solution is directly related to the maximum proportion of samples belonging to any class. We empirically validate our predictions and systematically analyze the effects of shedding the impractical assumptions one at a time, as well as incorporating nonlinearities. We observe that our predictions are considerably robust in most cases, allowing us to extend our conclusions to more practical learning setups.

Comments15 pages, 8 figures; Accepted at HiLD@ICML'26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑