arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经逼近与极小极大回归中网络规模与参数幅度的最优权衡

Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression

Baicheng Li, Zuowei Shen, Haizhao Yang, Shijun Zhang

arXiv 2609.25710首次发表:更新:

发表机构

University of Maryland; National University of Singapore; The Hong Kong Polytechnic University(马里兰大学; 新加坡国立大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对神经网络逼近与回归,在固定深度下建立宽度与参数幅度的最优权衡,证明二进三角激活网络达到最优逼近误差,并实现无对数损失的极小极大回归风险,同时给出固定规模下的最优参数配置。

AI 中文摘要

神经网络的统计精度取决于其逼近能力以及从数据拟合的类别复杂度。虽然增大网络规模是提升逼近效果的自然途径,但参数幅度提供了另一种资源,其作用必须在逼近和统计两个层面加以量化。我们使用一种基本的有界1-Lipschitz二进三角激活函数,在固定深度下建立了宽度与幅度的精确权衡。对于[0,1]^d上0<β≤1的单位β-Hölder球,当网络宽度满足N≥2d+3且参数幅度以T≥1为界时,对于0<p<∞,最优L^p逼近误差的阶为[N^2 log(eNT)]^{-β/d}。对于每个固定的全局Hölder激活函数,匹配的下界成立;其Hölder指数影响常数但不影响速率。在有界设计密度和独立中心次高斯噪声下,深度为23的完整裁剪类上的近似最小二乘达到了经典Hölder极小极大风险O(M^{-2β/(2β+d)}),当N^2 log(eNT)与M^{d/(2β+d)}同阶时无对数损失,其中M为样本量。这产生了从单位参数半径到固定网络规模的连续统计最优选择。在固定规模下,具有至多8d+7个非零参数的四隐藏层给出了近最优半径,而具有至多8d+27个非零参数的六层在逼近误差η下达到最优阶log T=O(η^{-d/β})。相同的解码方法也产生了固定规模的Transformer逼近。

英文摘要

The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the unit $β$-Hölder ball on $[0,1]^d$ with $0<β\leq1$, the optimal $L^p$ approximation error for $0<p<\infty$ is of order $[N^2\log(eNT)]^{-β/d}$ when the network width satisfies $N\geq2d+3$ and the parameter magnitudes are bounded by $T\geq1$. Matching lower bounds hold for every fixed globally Hölder activation; its Hölder exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical Hölder minimax risk $\mathcal{O}(M^{-\frac{2β}{2β+d}})$ without logarithmic loss whenever $N^2\log(eNT)\asymp M^{\frac{d}{2β+d}}$, where $M$ is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order $\log T=\mathcal{O}(η^{-d/β})$ at approximation error $η$. The same decoding method also yields fixed-size Transformer approximation.

Comments71 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑