AI 中文总结
针对深度神经网络缺乏闭式解与凸优化目标的问题,提出度量空间上未知Lipschitz函数的两阶段闭式重构公式$\hat{f}$,给出高概率一致恢复保证,证明其在函数空间、参数空间及前向传播三方面的最优性,可由稀疏ReLU网络精确实现。
AI 中文摘要
核岭回归(KRR)、支持向量回归(SVR)等经典机器学习方法在计算与分析层面均具备可处理性,其估计量要么存在闭式表达式,要么通过最小化凸训练目标得到;而深度神经网络通常不具备这两种特性。针对该问题,本文提出一种简洁的“两阶段”组合闭式公式$\hat{f}$,用于基于$N$个独立同分布含噪观测值,重构度量空间$(\mathcal X,ρ)$上的未知Lipschitz函数$f:\mathcal{X}\to \mathbb{R}$。\n本文核心结果为高概率一致($L^{\infty}$)恢复保证,可同时控制逼近误差与统计误差,且优化误差为零;特别地,本文未假设可通过预言机访问近似经验风险最小化(ERM)结果。次要核心结果从三个互补维度证明了所提公式的最优性:1)函数空间层面:在阿尔福斯正则(Ahlfors-regular)度量空间上,由该公式参数化的假设类达到最优的胖粉碎维数;2)参数空间层面:其对参数的依赖具有最大数值稳定性,即无法通过降低模型参数的Lipschitz依赖度来获得更小的逼近误差;3)前向传播层面:其对输入的依赖具有最高正则性,与目标函数$f$的Lipschitz常数相匹配。当$\mathcal X=[0,1]^d$配备$\ell^\infty$范数时,$\hat{f}$可通过深度为$\mathcal{O}(\log(N))$、含$\mathcal{O}(N)$个非零参数的ReLU多层感知机(ReLU-MLP)算法实现,也可通过精确的ReLU多头Transformer实现。
英文摘要
Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks. We address this by introducing a simple closed-form ``two-stage'' compositional formula $\hat{f}$ for reconstructing an unknown Lipschitz function $f:\mathcal{X}\to \mathbb{R}$ on a metric space $(\mathcal X,ρ)$ from $N$ i.i.d. noisy observations. Our main result is a high-probability uniform ($L^{\infty}$) recovery guarantee that jointly controls approximation and statistical errors while enjoying an optimization error of zero; in particular, we do not assume oracle access to an approximate ERM. Our secondary main results establish the optimality of our formula in three complementary senses. 1) Function space: On Ahlfors-regular metric spaces, the hypothesis class parameterized by our formula attains the optimal fat-shattering dimension. 2) Parameter space: Its dependence on the parameters is maximally numerically stable, in the sense that a smaller approximation error cannot be achieved with a smaller Lipschitz dependence on the model parameters. 3) Forward pass: Its dependence on the input is maximally regular, matching the Lipschitz constant of the target function $f$. When $\mathcal X=[0,1]^d$ is equipped with the $\ell^\infty$ norm, $\hat{f}$ admits algorithmic ReLU-MLP and exact ReLU-multi-head transformer realizations of depth $\mathcal{O}(\log(N))$ with $\mathcal{O}(N)$ nonzero parameters.
Comments75 pages, 17 figures