arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00420stat.MLcs.LG

可迁移的图元网络

Transferable Graph Metanetworks

Yuxin Ma, Adir Dayan, Yam Eitan, Haggai Maron, Soledad Villar

首次发表
浏览论文内容

中文总结 AI 辅助

提出可迁移图元网络,通过不变性和连续性修改,使元网络性能可跨不同宽度网络迁移,在μP参数化下实现高达42倍宽度泛化,并给出理论保证。

中文摘要 AI 辅助

权重空间网络(或称元网络)以另一个神经网络的权重作为输入,并预测其属性。以往的大多数工作都是在固定大小或少数几个固定大小的输入网络上训练此类模型,并在分布内进行评估。少数尝试进行分布外大小泛化的研究仍然范围有限,且仅取得了有限的成功。因此,在小网络上训练并在大得多的网络上进行评估的潜在效率提升在很大程度上尚未实现。我们提出了可迁移的图元网络,它通过一系列修改扩展了图元网络范式,使得性能能够在不同宽度的输入网络之间迁移。这些修改遵循两个原则:对不同宽度的网络表示相同函数的方式具有不变性,以及连续性,即表示相似函数的权重能获得相似的预测。我们进一步研究了对于从随机初始化独立训练的输入网络,大小泛化是否可能实现。实验上,我们的修改在我们考虑的每个任务上都显著提高了大小泛化性能。在最大更新参数化($\mu$P)下训练的输入网络上性能最强,其鲁棒性可保持到训练宽度的$42\times$。理论上,我们用无限宽度极限理论解释了这些观察:我们证明了我们的模型在$\mu$P训练输入上的大小泛化保证,并解释了为什么在其他参数化下可能会失败。

英文摘要

A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such models on input networks of one or a few fixed sizes and evaluates them in-distribution. The few attempts at out-of-distribution size generalization remain limited in scope and have achieved only modest success. Consequently, the potential efficiency gains of training on small networks and evaluating on much larger ones remain largely unrealized. We propose Transferable Graph Metanetworks, which extend the graph metanetwork paradigm with a set of modifications that make performance transferable across input networks of different widths. The modifications follow two principles: invariance to the ways in which networks of different widths represent the same function, and continuity, such that weights representing similar functions receive similar predictions. We further study whether size generalization is possible for input networks trained independently from random initialization. Empirically, our modifications significantly improve size generalization on every task we consider. Performance is strongest on input networks trained under the maximal-update parameterization ($μ$P), where it remains robust up to $42\times$ the training width. Theoretically, we explain these observations with infinite-width limit theory: we prove size-generalization guarantees for our model on $μ$P-trained inputs, and explain why it can fail under other parameterizations.

↑