arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有内部宽度为一的残差神经网络通用逼近的最小块宽度

Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One

Qi Zhou, Xuan Zhou, Xiao-Song Yang

arXiv 2607.04597首次发表:更新:

发表机构

School of Mathematics and Statistics, Huazhong University of Science and Technology; Hubei Key Laboratory of Engineering Modeling and Scientific Computing, Huazhong University of Science and Technology(华中科技大学数学与统计学院; 湖北省工程建模与科学计算重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究残差神经网络通用逼近性质,针对特定激活函数建立块宽度上下界,证明内部宽度为1时在紧致域上实现\(L^p\)逼近的最小块宽度,还给出特定块宽度网络能实现一致逼近及宽度小于某值不能逼近所有目标函数的结论。

AI 中文摘要

在本文中,我们研究了残差神经网络的通用逼近性质,并获得了一些新结果。对于输入和输出维度\(d_x\)和\(d_y\),以及LeakyReLU、ReLU、类ReLU激活函数,建立了块宽度的上下界。为了在任何紧致域上实现\(L^p\)逼近\((1\leq p <+\infty)\),我们表明当内部宽度为1时,精确的最小块宽度是\(\max\{d_x,d_y\}\)。此外,我们表明在每个残差分支的内部宽度为1的约束下,具有块宽度\(\min\{d_x + d_y, \max\{2d_x + 1, d_y\}\}\)的残差神经网络可以在任何紧致域上实现一致逼近。此外,对于任何激活函数族,我们证明无论内部宽度如何,块宽度小于\(\max\{d_x, d_y\}\)的残差神经网络在\(L^p\)意义和一致意义上都不能逼近所有目标函数。

英文摘要

In this paper, we study the universal approximation property of residual neural networks. For input and output dimensions $d_x$ and $d_y$, and LeakyReLU, ReLU, ReLU-like activation functions, the upper and lower bounds of the minimum block width are established. To achieve $L^p$ approximation $(1\leq p <+\infty)$ on any compact set, we show that the exact minimum block width is $\max\{d_x,d_y\}$ when each residual branch has inner width 1. Furthermore, we show that residual neural networks with block width $\min\{d_x+d_y, \max\{2d_x+1,d_y\}\}$ can achieve uniform approximation on any compact set under the constraint that each residual branch has inner width 1. Besides, for any activation function family, we prove that there exist functions that cannot be approximated by residual neural networks with block width less than $\max\{d_x, d_y\}$, both in the $L^p$ sense and the uniform sense, regardless of inner width. Consequently, for LeakyReLU, ReLU, ReLU-like activation functions and $d_y\geq 2d_x+1$, the exact minimum block width for uniform approximation is $d_y$ when each residual branch has inner width 1.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑