AI 中文总结
该研究针对带Xavier初始化的宽随机tanh神经网络,通过分离最高阶导数线性项等方法,得到与深度无关的一阶导数界及多项式增长的混合导数界,还推导了欧几里得Lipschitz常数等相关高概率界。
AI 中文摘要
我们针对激活导数满足阶乘增长界的宽随机神经网络,建立了其混合输入导数的高概率界。我们的主要结果将这些估计专门应用于带有Xavier初始化的tanh网络。基于权重矩阵的欧几里得算子范数的直接确定性分析会产生通常随深度指数增长的导数界。我们表明,对于足够宽的高斯网络,通过分离最高阶导数的线性项并通过可测有限网控制相应的切方向,这种增长可以得到显著改善。对于带有高斯权重和Xavier初始化的标量输出tanh网络,我们证明存在常数C、C₀、C₁>0,使得当公共隐藏宽度满足n≥C(L³n₀²(1+log n₀)+L²(1+log(L/η)))时,以至少1-η的概率,对于每个非空u⊆[n₀]和每个x∈[0,1]^n₀,估计|Dᵘℛ_Φ⁽ᴸ⁾(x)|≤C₀|u|!(C₁L)^(|u|-1)∏_{j∈u}βⱼ(η,n₀)同时成立。因此,一阶导数界与深度无关,而阶数为|u|的无平方混合导数除坐标因子外最多随L^(|u|-1)多项式增长。作为推论,我们获得了网络实现的欧几里得Lipschitz常数和加权Sobolev范数的高概率界,后者将导数估计与拟蒙特卡罗(QMC)积分联系起来,并表明这种正则性如何能进入基于QMC的训练分析。
英文摘要
We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. Our main result specializes these estimates to $\tanh$ networks with Xavier initialization. A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that generally grow exponentially with the depth. We show that this growth can be substantially improved for sufficiently wide Gaussian networks by isolating the term that is linear in the highest-order derivative and controlling the corresponding tangent directions by measurable finite nets. For scalar-output $\tanh$ networks with Gaussian weights and Xavier initialization, we prove that there exist constants $C,C_0,C_1>0$ such that, whenever the common hidden width satisfies $n \geq C\left(L^3n_0^2(1+\log n_0)+L^2\left(1+\log(L/η)\right)\right)$, then, with probability at least $1-η$, the estimate $\left|D^u\mathcal{R}_{Φ^{(L)}}(x)\right| \leq C_0 |u|! (C_1L)^{|u|-1}\prod_{j\in u}β_j(η,n_0)$ holds simultaneously for every non-empty $u\subseteq[n_0]$ and every $x\in[0,1]^{n_0}$. Thus, the first-order derivative bound is independent of the depth, while a square-free mixed derivative of order $|u|$ grows at most polynomially as $L^{|u|-1}$, apart from the coordinate factors. As consequences, we obtain high-probability bounds for the Euclidean Lipschitz constant and for weighted Sobolev norms of the network realization. The latter connect the derivative estimates to quasi-Monte Carlo integration and indicate how such regularity can enter the analysis of QMC-based training.
Comments31 pages, 0 figures