arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

损失景观的隧穿:通过蒙特卡洛参数交换绕过记忆化

Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang

arXiv 2608.01833首次发表:更新:

发表机构

City University of Hong Kong; University of Pennsylvania; Air Liquide(香港城市大学; 宾夕法尼亚大学; 液化空气集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究受统计物理学启发,提出SAM-Swap优化插件,通过表征训练动力学的参数迁移率等指标,揭示标准优化的玻璃动力学特征,证明随机探索可加速grokking现象中的泛化。

AI 中文摘要

Grokking是神经网络训练中一种引人注目的现象,模型在经历长时间的纯记忆阶段后会突然实现泛化。尽管此前的研究试图通过权重范数等经典机器学习机制来解释它,但近期研究从统计物理学中获得类比,将grokking视为一种计算玻璃弛豫形式。该理论将初始记忆定义为“快速冷却”的结果,此时训练损失下降过快,形成玻璃态,随后进入“缓慢弛豫”阶段以实现最终泛化。尽管这一视角为代表性grokking理论提供了统一框架,但在很大程度上仍停留在宏观理论层面,未对训练动力学进行直接实证验证。本文引入三组件框架,通过参数迁移率(PM)以及玻璃动力学的两个代表性测量量:复本关联(RC)和分形维数(FD)来直接表征训练动力学。我们证明,标准优化方法呈现出清晰的玻璃动力学特征,且会将grokking网络固有地捕获在动力学受阻的记忆状态中,该状态具有坍缩的迁移率、强历史依赖性和通道式运动。这种定量一致性促使我们引入状态感知蒙特卡洛参数交换(SAM-Swap),这是一种受玻璃动力学中广泛使用的交换蒙特卡洛算法启发的优化插件,可加速泛化。通过比较SAM-Swap、权重衰减和高斯梯度噪声,我们发现加速泛化始终与参数空间中的随机探索相关,类似于物理学中的扩散。

英文摘要

Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statistical physics, framing grokking as a form of computational glass relaxation. This theory defines the initial memorization as a result of `fast cooling' where the training loss is reduced so quickly that a glass state is formed, followed by a `slow relaxation' towards final generalization. Although providing a unifying framework for representative grokking theories, this perspective has remained largely at the theoretical on macroscopic level without direct empirical validation on training dynamics. Here we introduce a three-component framework to directly characterize the training dynamics via parameter mobility (PM), and two representative measurements from glassy dynamics: replica correlation (RC) and fractal dimension (FD). We demonstrate that standard optimization presents clear signatures of glass dynamics and inherently traps the grokking network in a kinetic arrested memorization state with a collapsed mobility, strong history dependence, and channel-like motions. This quantitative agreement motivates us to introduce State-Aware Monte Carlo Parameter Swapping (SAM-Swap), an optimization plug-in that can accelerate generalization, inspired by swap Monte Carlo algorithm widely used in glass dynamics. Comparing SAM-Swap, weight decay, and Gaussian gradient noise, we find that accelerated generalization is consistently associated with random exploration in the parameter space, similar to diffusion in physics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑