arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过量级势视角理解神经网络的泛化

Generalization in Neural Networks Through the Lens of Magnitude Potential

Sahel Torkamani, Henry Gouk, Rik Sarkar

arXiv 2610.01633首次发表:更新:

发表机构

University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于度量量级理论的量级势概念,用于分析神经网络泛化,其比率与记忆化分数相关,能检测决策边界变化并提供顿悟的几何指标,且不受神经坍缩抑制影响。

AI 中文摘要

解释神经网络的泛化和训练动态仍然是一个挑战,人们已经开发了各种方法来研究这些现象的不同方面。在本文中,我们引入了量级势(magnitude potential)的概念——一个基于度量量级理论的量,它反映了一个任意点被给定集合表示得有多好。我们发现这个基本量可以用于检查神经泛化中的各种特征。在logit层计算的相对于某一类的量级势与相对于整个数据集的量级势之比,能够提供该点表示的信息。在实验中,单个训练点的这些比率被发现与Feldman记忆化分数相关。跨点聚合的量级势比率能够检测决策边界中的结构变化,并为模算术中的顿悟(grokking)提供几何指标。尽管量级势比率和神经坍缩(neural collapse)都与类内和类间的几何结构密切相关,但即使在神经坍缩被明确抑制的情况下,量级势比率仍然保持信息量。

英文摘要

Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of {\em magnitude potential} -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑