LP-NAS:基于线性规划的神经架构搜索
LP-NAS: Linear Programming-based Neural Architecture Search
浏览论文内容
中文总结 AI 辅助
本文提出LP-NAS框架,将线性规划应用于可微分NAS,其变体在CIFAR等数据集上的搜索与评估阶段均优于DARTS及其多种变体,且架构可迁移至ImageNet。
中文摘要 AI 辅助
神经架构搜索(NAS)旨在自动化神经网络架构设计,减少对人类专业知识的依赖。在各类NAS方法中,可微分NAS相较于传统NAS方法因效率和精度优势备受关注。由于可微分NAS将架构搜索空间松弛为连续域,因此可将连续优化原理应用于NAS。本文提出基于线性规划的NAS(LP-NAS),这是一种适用于广泛连续搜索空间的、基于数学规划的可微分NAS框架。LP-NAS利用验证损失梯度和训练损失海森矩阵构建线性规划(LP),以计算架构更新方向,该方向在保持模型参数最优性的同时提升泛化能力。通过遵循该LP导出的下降方向,LP-NAS可高效探索架构搜索空间,实现更快速、更有效的架构优化。我们引入LP-NAS的两种计算高效变体,即S-LP-NAS和R-LP-NAS。将LP-NAS应用于可微分架构搜索(DARTS)的搜索空间,得到S-LP-DARTS和R-LP-DARTS两种算法变体。两种变体在早期搜索迭代中均比标准DARTS算法收敛更快、验证性能显著更高。在CIFAR-10和CIFAR-100上的大量实验表明,LP-DARTS在架构搜索和评估阶段均优于标准DARTS。此外,我们在CIFAR-10数据集上将该方法与多种DARTS变体(P-DARTS、PC-DARTS和STO-DARTS)进行比较,验证了其有效性。我们还通过在ImageNet数据集上的实验验证了所发现架构的可迁移性。
英文摘要
Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expertise. Among the various NAS methods, differentiable NAS has gained prominence due to its efficiency and accuracy compared to conventional NAS approaches. Since differentiable NAS relaxes the architecture search space into a continuous domain, it is possible to apply principles from continuous optimization to NAS. In this paper, we propose Linear Programming-based NAS (LP-NAS), a mathematical programming-based framework for differentiable NAS that is applicable to a wide range of continuous search spaces. LP-NAS formulates a linear program (LP) using the validation-loss gradient and the training-loss Hessian to compute an architecture update direction that improves generalization while preserving the optimality of the model parameters. By following this LP-derived descent direction, LP-NAS efficiently navigates the architecture search space, leading to faster and more effective architecture optimization. We introduce two computationally efficient variants of LP-NAS, namely S-LP-NAS and R-LP-NAS. Applying LP-NAS to the Differentiable Architecture Search (DARTS) search space results in two algorithmic variants, S-LP-DARTS and R-LP-DARTS. Both variants achieve faster convergence and significantly higher validation performance during the early search iterations than the standard DARTS algorithm. Extensive experiments on CIFAR-10 and CIFAR-100 show that LP-DARTS outperforms standard DARTS in both the architecture search and evaluation phases. Additionally, we compare our approach with several DARTS variants (P-DARTS, PC-DARTS, and STO-DARTS) on the CIFAR-10 dataset and demonstrate its effectiveness. Furthermore, we validate the transferability of the discovered architectures through experiments on the ImageNet dataset.
发表机构
- IIT Kanpur(印度理工学院坎普尔分校)
- IIM Ahmedabad(印度管理研究所艾哈迈达巴德分校)
机构由 AI 辅助整理,请以论文原文为准。