发表机构
University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于度量量级理论的量级势概念,用于分析神经网络泛化,其比率与记忆化分数相关,能检测决策边界变化并提供顿悟的几何指标,且不受神经坍缩抑制影响。
AI 中文摘要
解释神经网络的泛化和训练动态仍然是一个挑战,人们已经开发了各种方法来研究这些现象的不同方面。在本文中,我们引入了量级势(magnitude potential)的概念——一个基于度量量级理论的量,它反映了一个任意点被给定集合表示得有多好。我们发现这个基本量可以用于检查神经泛化中的各种特征。在logit层计算的相对于某一类的量级势与相对于整个数据集的量级势之比,能够提供该点表示的信息。在实验中,单个训练点的这些比率被发现与Feldman记忆化分数相关。跨点聚合的量级势比率能够检测决策边界中的结构变化,并为模算术中的顿悟(grokking)提供几何指标。尽管量级势比率和神经坍缩(neural collapse)都与类内和类间的几何结构密切相关,但即使在神经坍缩被明确抑制的情况下,量级势比率仍然保持信息量。
英文摘要
Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of {\em magnitude potential} -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.