Grokking是否是插值流形的法双曲性丧失?
Is Grokking a Loss of Normal Hyperbolicity of the Interpolation Manifold?
浏览论文内容
中文总结 AI 辅助
该研究针对Grokking的泛化转变,通过残差雅可比矩阵的最小非零奇异值σ_min⁺(J)诊断,发现其在转变时未崩溃,为反对分岔假说、支持平滑收缩图景提供了初步证据。
中文摘要 AI 辅助
近期一系列研究将Grokking的记忆后阶段重新表述为约束优化:一旦网络对训练集进行插值,权重衰减会驱动沿零损失流形向更低范数的缓慢漂移。以动力系统的语言来说,这是一个快慢系统,其中插值流形扮演慢流形的角色。我们提出一个该框架自然引出但现有文献未解决的问题:泛化的急剧转变是否是该流形法双曲性的丧失,即类似折叠或分岔的事件,其中法向恢复方向变平?还是流形保持一致吸引,而泛化通过平滑漂移发生?我们提出一种简单、与优化器无关的诊断方法:残差雅可比矩阵的最小非零奇异值σ_min⁺(J),对于平方损失,它等于流形最慢的法向恢复速率。在训练以通过平方损失对模加法进行Grokking的两层ReLU网络上,σ_min⁺(J)在转变时不会崩溃;它仅在记忆前接近零,而在转变期间达到最大值。该结果在5个随机种子上成立,且最小的6个奇异值表现相同;也不存在子空间局部崩溃。这是反对分岔假说、支持平滑收缩图景的初步证据。我们明确指出,在Adam优化器下的单一设置、渐进转变实验并未证明分岔不存在;它仅限制了分岔可能隐藏的位置。
英文摘要
A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm. In the language of dynamical systems, this is a fast-slow system in which the interpolation manifold plays the role of a slow manifold. We ask a question that this framing makes natural but the existing literature does not address: is the sharp generalization transition a loss of normal hyperbolicity of that manifold: a fold- or bifurcation-like event in which a normal restoring direction goes flat? Or does the manifold stay uniformly attracting while generalization happens by smooth drift? We propose a simple, optimizer-agnostic diagnostic: the smallest nonzero singular value $σ_{\min}^{+}(\mathbf J)$ of the residual Jacobian, which, for the squared loss, equals the slowest normal restoring rate of the manifold. On a two-layer ReLU network trained to grok modular addition under squared loss, $σ_{\min}^{+}(\mathbf J)$ does not collapse at the transition; it is near zero only before memorization and attains its largest values during the transition. The result holds across five seeds, and the six smallest singular values behave identically; there is no subspace-local collapse either. This is preliminary evidence against the bifurcation hypothesis and in favor of the smooth-contraction picture. We are explicit that a single-setting, gradual-transition experiment under Adam optimizer does not prove the absence of a bifurcation; it constrains where one could hide.
发表机构
- Technische Universität Braunschweig(布伦瑞克工业大学)
机构由 AI 辅助整理,请以论文原文为准。