发表机构
RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究宽神经网络中单层与分层学习的差异,通过研究三层有限但大量隐藏单元的神经网络,发现训练输入到隐藏层权重泛化误差更小,且固定权重设置在参数空间有奇异性,揭示奇异性在宽神经网络中有重要作用。
AI 中文摘要
分层神经网络在人工智能中广泛应用,但其数学特性尚未完全明晰。在无限宽度极限下,已提出两种不同理论框架。本文研究具有有限但大量隐藏单元的三层神经网络,表明训练输入到隐藏层权重比固定它们产生更小泛化误差,且后者在参数空间呈现奇异性而前者没有,这表明奇异性在宽神经网络中也起着重要作用。
英文摘要
Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infinite-width limit, two different theoretical frameworks have been proposed. One reduces deep learning to kernel regression with a fixed kernel by assuming that the parameters remain close to their initialization, whereas the other allows the parameters to move away from their initialization, requiring the kernel itself to be optimized. In this paper, we study a three-layer neural network with a finite but large number of hidden units. We show that training the input-to-hidden weights yields a smaller generalization error than keeping them fixed. Furthermore, the latter setting exhibits singularities in the parameter space, whereas the former does not. These findings indicate that singularities play an essential role even in wide neural networks.