NGN:将神经网络规模学习为可微计数
NGN: Learning Neural Network Size as a Differentiable Count
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出神经发生网络(NGN),通过可微边界学习有序结构组件的数量,实现结构容量作为可微计数优化,并在多种架构上验证了与固定规模模型相当的性能。
AI中文摘要:
神经网络规模通常在训练之前选定,将架构选择与权重优化分离。我们引入了神经发生网络(NGN),这是一种可微参数化方法,用于学习模型应使用多少个有序结构组件。对于每个有序组件组,一个可学习的边界选择活跃前缀,同时训练模型参数。该边界可以从紧凑初始化开始增长,并且可以通过丢弃超出学习边界的组件来进行部署。受控实验考察了学习边界的收敛性、已部署前缀的性能,以及与固定规模模型和学习容量的替代方法的比较。然后,我们将相同的机制应用于多层感知机(MLP)、卷积网络、图网络、Transformer、状态空间模型、LoRA和适配器。在这些设置中,仅部署学习到的前缀通常对性能影响很小,并且所选架构的表现与在相同规模下训练的固定模型相似。这些结果表明,结构容量可以直接作为计数进行优化。
英文摘要:
Neural network size is usually chosen before training, separating architecture selection from weight optimization. We introduce the Neurogenesis Network (NGN), a differentiable parameterization for learning how many ordered structural components a model should use. For each ordered component group, one learnable boundary selects an active prefix while the model parameters are trained. The boundary can grow from a compact initialization and can be deployed by discarding components beyond the learned boundary. Controlled experiments examine convergence of the learned boundary, the performance of deployed prefixes, and comparisons with fixed-size models and alternative approaches to learning capacity. We then apply the same mechanism to MLPs, convolutional and graph networks, Transformers, state-space models, LoRA, and adapters. Across these settings, deploying only the learned prefix usually changes performance little, and the selected architectures perform similarly to fixed models trained at the same size. These results show that structural capacity can be optimized directly as a count.