对称性与奇异性
Symmetries and Singularities
浏览论文内容
中文总结 AI 辅助
本研究利用图结构与注意力参数的对称性,在教师-学生框架下解析估计图注意力模型的局部学习系数,使大规模LLC计算更可行。
中文摘要 AI 辅助
深度神经网络高度过参数化,不同的参数值可能表示相同的预测函数。这使得仅通过参数数量或Hessian矩阵的秩来衡量其有效复杂度变得困难。奇异学习理论通过局部学习系数(LLC)来解决这一问题,该系数刻画了模型在给定解附近的有效复杂度。现有的LLC估计方法通常依赖于后验采样,这对于大型神经网络而言计算成本高昂,使得在大规模场景下准确估计LLC变得困难。在本工作中,我们利用模型中的已知结构来简化分析,使LLC估计更加可行。具体而言,我们通过利用图结构和注意力参数中的对称性,研究图注意力模型的LLC。我们开发了一个通过教师-学生设置的解析框架,并在考虑对称性引起的简并性后,给出了显式的LLC估计。
英文摘要
Deep neural networks are highly over-parameterized, and different parameter values represent the same predictive function. This makes their effective complexity difficult to measure using only the number of parameters or the rank of the Hessian. Singular Learning Theory addresses this issue through the local learning coefficient (LLC), which characterizes the effective complexity of a model near a given solution. Existing methods for estimating the LLC often rely on posterior sampling, which can be computationally expensive for large neural networks. This makes accurate LLC estimation difficult at scale. In this work, we use known structures in the model to simplify the analysis and make LLC estimation more tractable. Specifically, we study the LLC of a graph attention model by exploiting symmetries in both the graph structure and the attention parameters. An analytic framework through a teacher--student setting, and explicit LLC estimates after considering the symmetry--induced degeneracies are developed.
发表机构
- Truth Audit Labs(真相审计实验室)
机构由 AI 辅助整理,请以论文原文为准。