arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过跳跃连接去除平面特征的伪局部极小值

Removing spurious minima for planar features by skip connections

Jakob Paul Zimmermann, Andrei Balakin, Moritz Grillo, Georg Loho

arXiv 2610.01728首次发表:更新:

发表机构

Freie Universität Berlin; Technische Universität Berlin; Max Planck Institute for Mathematics in the Sciences, Leipzig(柏林自由大学; 柏林工业大学; 马克斯·普朗克数学科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究证明在教师-学生设置中,可学习的线性跳跃连接能消除浅层ReLU网络的伪局部极小值,即使过参数化也无法移除的伪极小值可被跳跃连接去除,且学生特征位于教师特征张成空间内。

AI 中文摘要

理解损失景观是解释神经网络训练的核心,然而即使在简单模型中,其结构也仅被部分理解。我们研究了教师-学生设置中浅层、无偏置ReLU网络的高斯总体损失。这为研究特征学习和过参数化等基本方面提供了一个简单模型。对于具有正输出权重和平面特征的教师网络,我们证明一旦学生网络至少与教师网络一样宽,包含一个可学习的线性跳跃连接将消除所有具有非负学生输出权重的伪局部极小值。相反,在没有跳跃连接的情况下,我们构造了一个具有正输出权重且输入维度为二、仅三个隐藏神经元的固定教师网络,其伪局部极小值在任意学生宽度至少为三时持续存在。因此,可学习的线性跳跃连接可以消除在任意过参数化下持续存在的伪极小值。此外,我们证明了具有正输出权重的学生网络总是学习教师特征所张成的子空间:具有非负学生输出权重的局部极小值处的学生特征位于教师特征的张成空间中。对于二维ReLU网络,即使严重过参数化的学生网络的有效宽度也受教师宽度控制:具有正学生输出权重的每个临界点至多有教师神经元数量的两倍的不同学生特征方向。最后,我们将良性结果转移到任意规定半径的参数球上的经验极小值,所需的采样精度取决于该半径。

英文摘要

Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.

Comments43 pages, 4 figures. Under review. Accompanying Lean 4 formalization available at https://github.com/JayPiZimmermann/Removing-spurious-minima-for-planar-features-by-skip-connections

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑