arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25743math.STstat.TH

采用带剪枝的梯度下降学习深度神经网络回归估计

Learning of deep neural network regression estimates using gradient descent with pruning

Michael Kohler, Vincent Molinero Römer, Adam Krzyżak

AI总结:

本文提出采用带剪枝的梯度下降学习深度神经网络回归估计的方法,在回归函数光滑或预测变量集中于低维流形时可规避维数灾难,达到最优收敛速率,有限样本性能经模拟数据验证。

AI中文摘要:

本文考虑从独立同分布数据中估计回归函数,采用关于设计变量积分的L₂误差作为误差准则。将初始随机剪枝的全连接深度神经网络(以logistic激活函数作为压缩器),通过梯度下降拟合到数据,梯度下降过程中使用依赖于数据的非恒定步长选择。研究表明,当回归函数为(p,C)光滑时,该网络(至多差一个对数因子)达到最优极小极大收敛速率;若预测变量集中在低维流形附近,该估计可规避维数灾难。通过将其应用于模拟数据,展示了该估计的有限样本性能。

英文摘要:

Estimation of a regression function from independent and identically distributed data is considered. The $L_2$ error with integration with respect to the design variable is used as the error criterion. An initially randomly pruned fully connected deep neural network with logistic squasher as activation function is fitted to the data via gradient descent, using a data-dependent choice of non-constant stepsizes during gradient descent. It is shown that this network achieves (up to a logarithmic factor) the optimal minimax rate of convergence in case that the regression function is $(p,C)$--smooth. Here the estimate is able to circumvent the curse of dimensionality provided the predictors are concentrated in the neighborhood of a low dimensional manifold. The finite sample size performance of the estimate is illustrated by applying it to the simulated data.

↑