AI 中文总结
本文确定了浅层ReLU神经网络在临界Besov类上的最优逼近率,给出了尖锐的代数指数及两侧估计,并揭示了与有限元方法的比较优势。
AI 中文摘要
设$\mathbb D$为由$\operatorname{ReLU}^k$在有界Lipschitz域$\Omega\subset\mathbb R^d$上生成的归一化脊字典。我们确定了当外层$\ell^1$系数预算与$n$无关时,从$\mathbb D$进行有限$n$项逼近的尖锐代数速率。更精确地,设$d\ge3$,$k\in\mathbb N_+$,$0\le m\le k$,$0<p<1$,且$0<q\le1$。对于临界Besov空间$B_{p,q}^{k+d/p}(\Omega)$的单位球,在$H^m(\Omega)$度量下,最优代数指数为\\[ \min\left\{\frac{k-m+d/2}{d-1}, \frac{k-m+d/2+1/p-1/2}{d}\right\}. \\] 我们在此最小值两个分支确定的三个区域中证明了两侧估计,并给出了在过渡点及其之外的相应对数因子。在严格角区域以及当$q\le[1/2+(k-m+d/2)/(d-1)]^{-1}$时的过渡点,估计没有对数损失而一致。上界估计结合了临界小波稀疏性、局部Fourier-Radon表示以及方向和偏置的稳定分配。下界估计源于两个不同的障碍,即方向球上的径向脊逼近和联合方向-偏置空间中的Gevrey局部化。临界光滑性恰好是Besov正则性提供一致浅层网络变差界而无需额外尺度衰减的端点。在参数计数比较下,所得表示指数严格大于固定次数各向同性有限元的次数受限指数。最后这一陈述涉及最佳逼近,而非训练或计算复杂度。
英文摘要
Let $\mathbb D$ be the normalized ridge dictionary generated by $\operatorname{ReLU}^k$ on a bounded Lipschitz domain $Ω\subset\mathbb R^d$. We determine the sharp algebraic rate of finite $n$-term approximation from $\mathbb D$ when the outer $\ell^1$ coefficient budget is independent of $n$. More precisely, let $d\ge3$, $k\in\mathbb N_+$, $0\le m\le k$, $0<p<1$, and $0<q\le1$. For the unit ball of the critical Besov space $B_{p,q}^{k+d/p}(Ω)$, measured in $H^m(Ω)$, the optimal algebraic exponent is \[ \min\left\{\frac{k-m+d/2}{d-1}, \frac{k-m+d/2+1/p-1/2}{d}\right\}. \] We prove two-sided estimates in the three regimes determined by the two branches of this minimum and give the corresponding logarithmic factors at and beyond the transition. The estimates coincide without logarithmic loss in the strict angular regime and at the transition when $q\le[1/2+(k-m+d/2)/(d-1)]^{-1}$. The upper estimate combines critical wavelet sparsity, localized Fourier--Radon representations, and a stable allocation of directions and biases. The lower estimates arise from two distinct obstructions, namely radial ridge approximation on the direction sphere and Gevrey localization in the joint direction--bias space. The critical smoothness is exactly the endpoint at which Besov regularity provides a uniform shallow-network variation bound without additional scale decay. The resulting representation exponent is strictly larger than the degree-limited exponent for fixed-degree isotropic finite elements under a parameter-count comparison. This last statement concerns best approximation, not training or computational complexity.