arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38453cs.LGmath.STstat.MLstat.TH

通过最小范数插值的视角理解Grokking

Grokking through the Lens of Minimum-Norm Interpolation

Gil Kur, Ileana Rugina, Clémentine Carla Juliette Dominé, Marco Mondelli

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出统计理论,揭示正则化几何与信号稀疏性如何控制插值附近的泛化,证明零一律并刻画泛化增益,实验验证了理论预测,并发现最小范数插值的统计不稳定性。

中文摘要 AI 辅助

Grokking现象表明,拟合训练数据与学习潜在信号可能发生在截然不同的阶段。然而,现有理论对于这种延迟泛化如何依赖于归纳偏置和信号结构提供的定量洞察有限。我们的工作通过发展一种统计理论来填补这一空白,该理论刻画了正则化几何和信号稀疏性如何控制插值附近的泛化。具体而言,我们关注高维回归的典型设置,并识别出稀疏促进正则化使精确插值比近似拟合更准确的机制。在强过参数化的无噪声问题中,我们证明了一个零一律泛化定律,并构造了一族凸范数,其插值器从全零预测器的平凡风险过渡到精确恢复,同时保持训练误差为0。此外,当特征维度和样本量成比例时,我们提供了沿ℓ_r正则化路径的训练和泛化误差的精确刻画。这进而使我们能够量化在插值附近保持的泛化增益:我们表明,随着范数更促进稀疏性以及目标更稀疏,该增益增加,在无噪声数据和ℓ_1正则化下达到泛化的急剧下降。在模算术上训练的线性对角网络和Transformer上的实验证明了我们理论预测的普遍性。最后,超越grokking,我们的工作揭示了最小范数插值中的统计不稳定性:正则化强度的小扰动可能导致截然不同的泛化,同时保持较小的训练误差。

英文摘要

Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical theory that characterizes how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero--one generalization law and construct a family of convex norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error equal to $0$. Furthermore, when feature dimension and sample size are proportional, we provide a precise characterization of training and generalization errors along $\ell_r$-regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and $\ell_1$ regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate the generality of our theoretical predictions. Finally, beyond grokking, our work reveals a statistical instability in minimum-norm interpolation: small perturbations in the regularization strength can lead to drastically different generalization, while preserving small training error.

发表机构

  • ETH Zürich(苏黎世联邦理工学院)
  • Institute of Science and Technology Austria(奥地利科学技术研究所)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑