arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24210cs.LG

基于L1范数的Rademacher复杂度的依赖数据的早停规则

A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm

Duy Hoang, Bastien Berret, Olivier Bruneau, Laurent Fribourg

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种基于L1范数Rademacher复杂度的分析框架,无需训练即可估计早停最优时间,可应用于线性模型及经线性探测方法适配的非线性神经网络。

中文摘要 AI 辅助

训练神经网络需要平衡拟合训练数据与在未见过的输入上实现鲁棒性能之间的权衡,这种能力通常被称为泛化能力,由训练集上的经验风险(“经验损失”)与数据分布上的期望风险(“泛化误差”)之间的差距决定。现有方法通常通过数值方式估计泛化误差,需要梯度下降训练和“早停”策略。在本研究中,我们引入了一种无需训练即可估计早停最优时间的分析框架。文献中也有若干此类分析估计的研究,但它们通常基于随机矩阵理论,且往往对数据分布或协方差矩阵的特征值分布做出假设。相比之下,我们的工作基于Rademacher复杂度(RC),无需此类概率假设。由于理论和数值原因,用L1范数而非L2范数来表示RC更具相关性。我们聚焦于线性模型和线性回归问题,不过得益于“线性探测”方法,我们的结果可成功应用于非线性神经网络,MNIST分类示例对此进行了说明。

英文摘要

Training neural networks requires balancing the trade-off between fitting the training data and achieving robust performance on unseen inputs. This ability, commonly referred to as generalizability, is determined by the gap between the empirical risk on the training set (``empirical loss'') and the expected risk over the data distribution (``generalization error''). Existing approaches typically estimate the generalization error numerically, requiring gradient descent training and an ``early stopping'' strategy. In this work, we introduce an analytic framework that estimates the optimal time of early stopping without the need for training. Several works in the literature also give such analytical estimations, but they are generally based on random matrix theory and often make assumptions on the distribution of the data or the eigenvalue distribution of the covariance matrix. In contrast, our work is based on Rademacher complexity (RC) without needing such probabilistic assumptions. For both theoretical and numerical reasons, it is more relevant to express RC with the L1- norm rather than with the L2-norm. We focus on the case of linear models and the problem of linear regression. Thanks to the ``linear probing'' method, our results can, however, be successfully applied to nonlinear neural networks, as illustrated in the classification MNIST example.

发表机构

  • Université Paris-Saclay(巴黎-萨克雷大学)
  • CNRS(法国国家科学研究中心)
  • ENS Paris-Saclay(巴黎-萨克雷高等师范学校)
  • Inria(法国国家信息与自动化研究所)
  • CIAMS(运动与健康科学研究所)
  • LURPA(巴黎-萨克雷大学自动控制与生产实验室)
  • LMF(力学与流体实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑