arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

宽神经网络中单层学习与分层学习的统计差异

A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks

Sumio Watanabe

arXiv 2607.23397首次发表:更新:

发表机构

RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究宽神经网络中单层与分层学习的差异,通过研究三层有限但大量隐藏单元的神经网络,发现训练输入到隐藏层权重泛化误差更小,且固定权重设置在参数空间有奇异性,揭示奇异性在宽神经网络中有重要作用。

AI 中文摘要

分层神经网络在人工智能中广泛应用,但其数学特性尚未完全明晰。在无限宽度极限下,已提出两种不同理论框架。本文研究具有有限但大量隐藏单元的三层神经网络,表明训练输入到隐藏层权重比固定它们产生更小泛化误差,且后者在参数空间呈现奇异性而前者没有,这表明奇异性在宽神经网络中也起着重要作用。

英文摘要

Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infinite-width limit, two different theoretical frameworks have been proposed. One reduces deep learning to kernel regression with a fixed kernel by assuming that the parameters remain close to their initialization, whereas the other allows the parameters to move away from their initialization, requiring the kernel itself to be optimized. In this paper, we study a three-layer neural network with a finite but large number of hidden units. We show that training the input-to-hidden weights yields a smaller generalization error than keeping them fixed. Furthermore, the latter setting exhibits singularities in the parameter space, whereas the former does not. These findings indicate that singularities play an essential role even in wide neural networks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑