arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrossGMN:用于跨架构权重空间变换的图元网络

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

Adir Dayan, Yam Eitan, Haggai Maron

arXiv 2610.01649首次发表:更新:

发表机构

Technion – Israel Institute of Technology; NVIDIA Research(以色列理工学院; 英伟达研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对跨架构权重空间变换,提出CrossGMN图元网络,通过对称保持的跨网络消息传递实现等变算子,具备普适性,在模型压缩中加速蒸馏最高8.89倍并支持跨数据集迁移。

AI 中文摘要

权重空间网络直接对其他神经网络的参数进行操作,能够实现预测模型属性、编辑已训练模型以及生成权重等任务。神经元置换等权重空间对称性使得等变性成为关键的设计原则。然而,现有的等变权重空间架构主要针对保持网络架构不变的变换进行研究。相比之下,许多实际变换(包括模型压缩和放大)将已训练的源网络映射到具有不同架构的目标网络。在此设置下,源网络和目标网络的置换对称性作用于不同的参数空间,使得等变性的表述不那么直接。我们解决这一不匹配问题的关键思路是重新表述具有两个输入的跨架构算子:一个已训练的源网络和一个目标网络的初始化。这使我们能够定义等变的跨架构算子,利用源网络的信息来优化目标网络的初始化,同时保持对源网络置换的不变性和对目标网络置换的等变性。基于这一表述,我们引入了CrossGMN,一种图元网络,通过保持对称性的跨网络消息传递来联合处理两个网络。我们证明了在一般位置假设下,CrossGMN对于紧致集上的连续跨架构算子具有普适性。我们评估了CrossGMN在模型压缩中的应用,即预测较小网络的参数以加速后续的知识蒸馏。在二维和三维隐式神经表示(INRs)以及使用多层感知机(MLPs)、卷积神经网络(CNNs)和视觉变换器(Vision Transformers)的图像分类任务中,CrossGMN将蒸馏速度提升了最高达8.89倍,能够跨数据集迁移而无需重新训练(3.78倍),并且单一模型可以加速从异构源架构到共同目标架构的压缩。

英文摘要

Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑