发现神经网络参数空间中的对称性
Discovering Symmetries in Neural Network Parameter Spaces
浏览论文内容
中文总结 AI 辅助
本文通过无穷小条件形式化数据相关参数对称性,提出联合学习群生成器与非线性作用映射的框架,并利用子网络扩展实现自动发现,成功在多种架构(含预训练Transformer)中识别出对称性。
中文摘要 AI 辅助
参数空间对称性对于理解神经网络的损失景观、训练动态和泛化能力至关重要。然而,系统地识别这些对称性仍然是一个挑战。在本文中,我们形式化了数据相关的参数对称性,并通过无穷小条件刻画了损失不变性和群作用公理,这些条件为联合学习群生成器和非线性作用映射提供了目标。我们的框架系统地揭示了参数对称性,包括先前未知的对称性。为了研究更大的网络,我们建立了子网络对称性扩展到完整模型的条件。同样的构造给出了一个显式的有限批次对称性族,既提供了分析示例,也为通过小子网络进行发现奠定了基础。利用无穷小刻画和子网络构造,我们实现了一个用于自动发现参数对称性的框架,并成功地在各种架构中发现了对称性,包括预训练的Transformer模型。
英文摘要
Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter symmetries and characterize loss invariance and the group-action axioms through infinitesimal conditions, which provide objectives for jointly learning group generators and nonlinear action maps. Our framework systematically uncovers parameter symmetries, including previously unknown ones. To study larger networks, we establish conditions under which subnetwork symmetries extend to the full model. The same construction gives an explicit family of finite-batch symmetries, providing both analytical examples and a foundation for discovery through small subnetworks. Using the infinitesimal characterization and subnetwork construction, we implement a framework for automated discovery of parameter symmetries, and successfully uncovered symmetries in various architectures, including pretrained transformer models.
发表机构
- Harvard University(哈佛大学)
- IBM Research(IBM研究院)
- Northeastern University(东北大学)
- University of California, San Diego(加利福尼亚大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。