arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更平坦的极小值是否能推动更好的泛化?Grokking中的算法分离

Do Flatter Minima Drive Better Generalization? An Algorithmic Separation in Grokking

Mohnish Harwani

arXiv 2610.11206首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对Grokking场景,发现仅用锐度感知最小化(SAM)无法可靠诱导泛化过渡,SAM结合权重衰减可将过渡加速4倍,且平坦性无法区分记忆与泛化解,为平坦性驱动泛化提供了可解释说明。

AI 中文摘要

长期以来,平坦的损失景观一直与神经网络更好的泛化能力相关联,但它作为泛化因果机制的作用尚未明确确立。Grokking为理解这种区别提供了独特的测试平台:模型倾向于使用非泛化结构拟合观测数据,并在该状态下持续很长时间,仅在特定训练条件下才会过渡到泛化状态。本研究探讨平坦的损失景观是否能成为这种过渡的驱动机制。尽管近期研究认为平坦性是该过渡的必要几何条件,但我们发现,使用锐度感知最小化(SAM)将训练偏向更平坦的解,虽能产生更平坦的解,却不足以可靠地诱导该过渡。然而,当SAM与权重衰减等驱动泛化的机制结合时,会出现一个有趣的特性:SAM可在周期级别上将向泛化解的过渡加速多达4倍。我们使用兼具记忆解和泛化解的最小插值两层ReLU模型,从理论上梳理了SAM与权重衰减之间的关系。结果表明,即使在这种简单设置中,仅靠平坦性也无法区分记忆解与泛化解,而权重衰减更倾向于泛化解。但在局部稳定性分析下,存在一个区间:记忆插值器在梯度下降下是局部稳定的,而在低范数区域的SAM下是不稳定的,这可解释SAM加速该过渡的能力。总体而言,我们的结果为平坦性在驱动泛化中的作用提供了更具可解释性的解释,尤其是在模型易通过学习非泛化结构最小化损失的场景中。

英文摘要

Flat loss landscapes have long been linked to better generalization in neural networks. However, its role as a causal mechanism for generalization is less established. Grokking provides an unique testbed to understand this distinction: models are prone to fit observed data using non-generalizing structure and remain in that regime for prolonged periods, transitioning to generalization only under particular training conditions. In this work, we study whether flat loss landscapes can act as a driving mechanism in this transition. While recent work has argued for flatness as a necessary geometric condition for this transition, we find that biasing training toward flatter solutions using sharpness-aware minimization (SAM) is insufficient to reliably induce this transition, despite producing flatter solutions. However, when SAM is paired with mechanisms that drive generalization such as weight decay, an interesting property emerges: SAM can accelerate the transition to generalizing solutions by up to 4x at the epoch-level. We theoretically untangle this relationship between SAM and weight decay using a minimal interpolating two-layer ReLU model with both memorizing and generalizing solutions. We show that even in this simple setup, flatness alone cannot distinguish a memorizing solution from a generalizing one, while weight decay favors generalizing solutions. However, under a local stability analysis, there exists a window where a memorizing interpolant is locally stable under gradient descent but unstable under SAM in the low-norm regime, which can explain SAM's ability to accelerate this transition. Overall, our results provide a more interpretable account of the role of flatness in driving generalization, especially in settings where models are vulnerable to minimizing loss through learning non-generalizing structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑