arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

浅神经网络的景观分析:三次激活和仿射目标函数临界点的完全分类

Landscape analysis for shallow neural networks: Complete classification of critical points for cubic activation and affine target functions

Shokhrukh Ibragimov, Ilkhom Mukhammadiev, Diyora Salimova

arXiv 2607.15173首次发表:更新:

AI 中文总结

研究浅多项式神经网络在仿射目标函数训练下的优化景观,给出全局极小值存在/不存在准则,对三次激活损失函数的临界点完全分类,明确不同情况下临界点性质及隐藏神经元贡献等。

AI 中文摘要

本文研究了具有\(\mathfrak{h} \in \mathbb{N}\)个隐藏层神经元、一维输入和输出层以及\(d \in \mathbb{N}\)次单项式激活的浅多项式神经网络(PNN)在非恒定仿射线性目标函数训练下的真实损失所诱导的优化景观。第一个主要结果为任意激活度\(d\)提供了关于全局极小值的尖锐存在/不存在准则及必要结构条件。表明损失的下确界总是零,且至少有\(d\)个活跃且可见的隐藏神经元(即内外权重非零的隐藏神经元)且枢轴两两不同时可达到。若\(\mathfrak{h} < d\),则下确界无法达到且任何参数极小化序列必然发散到无穷。第二个主要结果对三次激活的损失函数的所有临界点进行了完全分类。表明损失景观不存在局部极大值,临界点不能恰好有两个不同枢轴,全局极小值至少需要三个不同枢轴,无活跃隐藏神经元的临界点仅对应鞍点,非全局局部极小值和非平凡鞍点仅出现在所有枢轴重合的网络中。此外,非全局局部极小值要求所有隐藏神经元活跃且可见,且恰好有一个隐藏神经元的斜率符号与目标函数匹配。第二个主要结果还保证,不是全局极小值的临界点的每个隐藏神经元对其相应的实现函数要么有输入依赖贡献要么贡献为零,但没有非零的与输入无关的贡献。

英文摘要

In this paper, we study the optimization landscape induced by the true loss for shallow polynomial neural networks (PNNs) with $\mathfrak{h} \in \mathbb{N}$ neurons on the hidden layer, one-dimensional input and output layers, and a monomial activation of degree $d \in \mathbb{N}$, trained against a non-constant affine linear target function. Our first main result provides for arbitrary activation degree $d$ a sharp existence/non-existence criterion for \emph{global minimizers} with necessary structural conditions. We show that the infimum of the loss is always zero and achievable with at least $d$ active and visible hidden neurons -- that is, hidden neurons with non-zero inner and outer weights -- with pairwise distinct pivots. In contrast, if $\mathfrak{h} < d$, then the infimum cannot be attained and any minimizing sequence of parameters necessarily diverges to infinity. In the second main result, we provide a complete classification of all critical points of the loss function for the cubic activation. We show that the loss landscape admits no \emph{local maximizers}, critical points cannot have exactly two distinct pivots, global minimizers require at least three distinct pivots, critical points with no active hidden neurons correspond to \emph{saddle points} only, and consequently, \emph{non-global local minimizers} and non-trivial saddle points arise only in networks where all pivots coincide. Moreover, non-global local minimizers require all hidden neurons to be active and visible with exactly one hidden neuron having a slope sign matching that of the target function. Our second main result also guarantees that each hidden neuron of a critical point that is not a global minimizer has either input-dependent or zero contribution, but has no nonzero input-independent contribution, to its corresponding realization function.

Comments37 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑