具有可调深度和宽度的ReLU神经网络对解析函数的逼近
Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width
- Department of Mathematics, The University of Hong Kong(香港大学数学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究通过ReLU网络在$(N,L)$表征下对解析函数的逼近,推导了$\mathcal{O}\left(N^{-C L^{\tau}}\right)$的逼近率,揭示解析函数逼近中深度比宽度更关键,通过精细构造网络克服光滑度与精度权衡的技术难点。
AI中文摘要:
与大多数通过单个参数(如网络参数总数)来表征神经网络逼近结果的研究不同,[shen2020deep]率先将逼近率表征为宽度参数$N$和深度参数$L$的联合函数,从而赋予了更大的架构灵活性。现有使用$(N,L)$表征的工作聚焦于有限光滑度$s$的函数类,建立了典型逼近率$\mathcal{O}\left(N^{-2s/d}L^{-2s/d}\right)$,其中$d$表示输入维度,这表明网络深度和宽度对这些类起对称作用。相比之下,本文通过ReLU网络在$(N,L)$表征下建立了解析函数(具有无限光滑度)逼近的上界。具体而言,我们推导了$\mathcal{O}\left(N^{-C L^{\tau}}\right)$的逼近率,其中$C>0$是某个常数,$\tau>0$是受$L$与$N$关系影响的参数。特别地,当$N$大致按$L^d$缩放时,$\tau = 1$。我们的发现揭示了在解析函数逼近中深度比宽度起着更关键的作用。获得此类上界的主要技术难点在于光滑度参数与逼近精度之间的权衡。为克服这一困难,我们采用了几个ReLU网络的精细构造来逼近幂函数、多元乘法和多项式,这可能具有独立的研究价值。
英文摘要:
In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width parameter $N$ and the depth parameter $L$, thereby granting greater architectural flexibility. Existing works using the $(N,L)$-characterization focus on function classes with finite smoothness $s$, establishing a typical approximation rate of $\mathcal{O}\left(N^{-2s/d}L^{-2s/d}\right)$ with $d$ denoting the input dimension, which indicates that network depth and width play symmetric roles for these classes. In contrast, this paper establishes upper bounds for the approximation of analytic functions, which possess infinite smoothness, via ReLU networks under the $(N,L)$-characterization. Specifically, we derive approximation rates of $\mathcal{O}\left(N^{-C L^τ}\right)$, where $C>0$ is some constant and $τ>0$ is a parameter influenced by the relation between $L$ and $N$. In particular, $τ=1$ if $N$ scales roughly as $L^d$. Our findings reveal that depth plays a more critical role than width in the context of analytic function approximation. The main technical difficulty of obtaining such upper bounds lies in the trade-off between the smoothness parameters and the approximation accuracy. To overcome this difficulty, we employ refined constructions of several ReLU networks to approximate power functions, multivariate multiplication, and polynomials, which may be of independent interest.