发表机构
FAST National University of Computer & Emerging Sciences (NUCES)(FAST国家计算机与新兴科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出MFSPNet,将无模型代理预测器融入PSO框架,通过VLE-EMA和稠密连接策略实现高效NAS,在多数据集上以低计算成本取得具竞争力性能。
AI 中文摘要
神经网络架构搜索(NAS)已成为自动设计深度神经网络的强大范式,但其实际应用常受限于巨大的计算成本。为缓解完整训练评估的高昂开销,基于代理的方法被提出以高效估计网络性能。然而,现有方法——尤其是基于模型的代理——需要训练大量候选架构并涉及额外的优化开销。本研究提出一种用于演化卷积神经网络架构的无模型代理PSO网络(MFSPNet),该方法在粒子群优化(PSO)框架内集成了轻量级无模型代理预测器,无需预训练代理模型。具体而言,MFSPNet引入两项关键贡献:(1)验证损失驱动的指数移动平均估计器(VLE-EMA),用于捕捉早期泛化行为以实现可靠的架构排名;(2)基于块的稠密连接策略,可有效堆叠演化后的块同时缓解梯度消失问题,该设计还提升了所学块跨数据集的可迁移性。大量实验表明,MFSPNet以更低的计算成本取得了具竞争力的性能:在包含十次独立运行的一致训练协议下,该方法在CIFAR-10、CIFAR-100和SVHN上的错误率分别为3.91%、17.68%和1.91%,在ImageNet上的top-1/top-5错误率为28.29%/12.82%,且架构搜索所需GPU时间不足三天;受计算资源限制,ImageNet结果基于单次运行,仅作为可扩展性的指示。总体而言,MFSPNet为成本感知型神经网络架构搜索提供了高效可靠的框架。
英文摘要
Neural Architecture Search (NAS) has emerged as a powerful paradigm for automatically designing deep neural networks; however, its practical adoption is often limited by substantial computational cost. To alleviate expensive full-training evaluations, surrogate-based methods have been introduced to estimate network performance efficiently. Nevertheless, existing approaches-particularly model-based surrogates-require training many candidate architectures and involve additional optimization overhead. In this work, we propose a Model-Free Surrogate PSO Network (MFSPNet) for evolving convolutional neural network architectures. The proposed method integrates a lightweight model-free surrogate predictor within a particle swarm optimization (PSO) framework, eliminating the need for pre-trained surrogate models. Specifically, MFSPNet introduces two key contributions: (1) a validation-loss-driven exponential moving average estimator (VLE-EMA) that captures early generalization behavior for reliable architecture ranking; and (2) a block-based dense connection strategy that enables effective stacking of evolved blocks while mitigating vanishing-gradient issues. This design also facilitates transferability of learned blocks across datasets. Extensive experiments demonstrate that MFSPNet achieves competitive performance with reduced computational cost. Under a consistent training protocol with ten independent runs, the proposed method attains error rates of 3.91%, 17.68%, and 1.91% on CIFAR-10, CIFAR-100, and SVHN, respectively, along with top-1/top-5 error rates of 28.29%/12.82% on ImageNet, while requiring less than three GPU days for architecture search. Due to computational constraints, the ImageNet result is based on a single run and should be interpreted as indicative of scalability. Overall, MFSPNet provides an efficient and reliable framework for cost-aware neural architecture search.
Journal ref10.1016/j.eswa.2026.132442