arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14443cs.LGcs.AI

通过神经元门控与混合激活设计紧凑的神经架构

Designing Compact Neural Architectures via Neuron Gating and Mixed Activation

  • IIT Kanpur(印度理工学院坎普尔分校)
  • IIM Ahmedabad(印度管理学院艾哈迈达巴德分校)
  • Krishnamurthy Tandon School of AI(克里希纳穆尔蒂坦登人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

Abhishek Shukla, Ankur Sinha, Faiz Hamid

AI总结:

本研究提出基于神经元门控与混合激活的NAS方法,通过连续松弛优化架构空间,在MNIST、CIFAR-10上实现紧凑架构的高性能,优于DARTS,可优化过度参数化架构。

AI中文摘要:

神经架构搜索(NAS)天然可表述为一个双层优化问题:上层利用验证性能优化架构,下层利用训练损失训练网络参数。然而,NAS因离散架构决策、指数级增长的搜索空间及训练候选架构的高昂成本而计算成本高昂。本研究开发了一种适用于MLPs、CNNs、RNNs、Transformers等各类架构的通用NAS双层优化框架,以识别兼具强预测性能的紧凑架构。我们提出三种可扩展的公式,将离散的神经元级与激活级决策替换为连续松弛,从而可在原本为组合式的架构空间上进行可微优化。这些公式衍生出三种NAS方法:基于神经元门控的NAS(NAS-NG)、基于混合激活的NAS(NAS-MA),以及基于神经元门控与混合激活的NAS(NAS-NGMA)。在MNIST和CIFAR-10数据集上针对MLPs和CNNs开展的实验表明,所提方法始终能识别出兼具竞争力或更优预测性能的紧凑架构。在MNIST上,NAS-NGMA以769万个MLP参数达到98.68%的测试准确率,而NAS-NG仅用26万个CNN参数就达到99.63%的准确率;在CIFAR-10上,所提方法始终优于基准方法DARTS。进一步实验表明,NAS-NG可对过度参数化及文献最优架构进行优化,在提升准确率的同时减少参数数量。这些结果确立了松弛双层优化作为离散NAS的可扩展替代方案的地位,并为高效的神经元与激活级架构优化提供了通用框架。

英文摘要:

Neural Architecture Search (NAS) is naturally formulated as a bilevel optimization problem, where the upper-level optimizes the architecture using validation performance and the lower-level trains network parameters using training loss. However, NAS is computationally expensive due to discrete architectural decisions, exponentially growing search spaces, and the high cost of training candidate architectures. This work develops a general bilevel optimization framework for NAS across diverse architectures, including MLPs, CNNs, RNNs, and Transformers, to identify compact architectures with strong predictive performance. We propose three scalable formulations that replace discrete neuron- and activation-level decisions with continuous relaxations, enabling differentiable optimization over otherwise combinatorial architecture spaces. These formulations give rise to three NAS methods: NAS based on Neuron Gating (NAS-NG), NAS based on Mixed Activation (NAS-MA), and NAS based on Neuron Gating and Mixed Activation (NAS-NGMA). Experiments on MLPs and CNNs using MNIST and CIFAR-10 show that the proposed methods consistently identify compact architectures with competitive or improved predictive performance. On MNIST, NAS-NGMA achieves 98.68% test accuracy with 7.69M MLP parameters, while NAS-NG achieves 99.63% accuracy with only 0.26M CNN parameters. On CIFAR-10, the proposed methods consistently outperform vanilla DARTS. Further experiments demonstrate that NAS-NG can optimize substantially over-parameterized and literature-optimal architectures, improving accuracy while reducing parameters. These results establish relaxed bilevel optimization as a scalable alternative to discrete NAS and provide a general framework for efficient neuron- and activation-level architecture optimization.

补充信息

↑