AI 中文总结
研究神经网络‘顿悟’现象中表示几何与延迟泛化的关系,引入几何维度正则化方法,发现其能改变‘顿悟’动态,加速泛化,为研究和影响神经网络延迟泛化提供实用途径。
AI 中文摘要
‘顿悟’是一种现象,即神经网络最初记忆训练数据,只有在长时间优化后才展现出强大的泛化能力。尽管近期有广泛研究,但影响‘顿悟’出现及时间的因素仍未完全明晰。我们研究表示几何与延迟泛化之间的关系,发现维度坍缩在所有评估设置中都先于‘顿悟’出现。基于此,我们引入几何维度正则化(GeomDR),一种简单的谱正则化方法,在训练期间修改隐藏表示的有效维度。在模块化加法、模块化除法和置换合成任务中,GeomDR 一致地改变‘顿悟’动态,并可根据干预时间表和目标维度大幅加速泛化的开始。在几种设置中,相对于标准 AdamW 训练,‘顿悟’加速了多达 52 倍。在多层感知器和变压器中都观察到了类似的定性效果。这些结果共同表明,表示几何可作为‘顿悟’的有效控制信号,并证明几何干预为研究和影响神经网络中的延迟泛化提供了一种实用方法。
英文摘要
Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization. Despite extensive recent study, the factors influencing the emergence and timing of grokking remain incompletely understood. We investigate the relationship between representation geometry and delayed generalization. We find that dimensionality collapse consistently precedes the onset of grokking in all evaluated settings. Motivated by these observations, we introduce Geometric Dimensionality Regularization (GeomDR), a simple spectral regularizer that modifies the effective dimensionality of hidden representations during training. Across modular addition, modular division, and permutation composition tasks, GeomDR consistently alters grokking dynamics and can substantially accelerate the onset of generalization depending on the intervention schedule and target dimensionality. In several settings, grokking is accelerated by up to 52 times relative to standard AdamW training. Similar qualitative effects are observed in both multilayer perceptrons and transformers. Together, these results suggest that representation geometry can serve as an effective control signal for grokking and provide evidence that geometric interventions offer a practical approach for studying and influencing delayed generalization in neural networks.