发表机构
The University of Tokyo; Sakana AI; Waseda University(东京大学; Sakana人工智能; 早稻田大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LESS提出一种轻量级进化超网络搜索方法,结合公平硬路径预热与CMA-ES离散搜索,在几分钟内实现高精度架构搜索,显著降低计算成本并保持竞争力。
AI 中文摘要
低成本神经架构搜索(NAS)必须既探索高性能架构,又可靠地识别它们,然而降低评估成本往往会削弱候选比较的保真度。免训练方法通过用初始化时测量的代理信号替代学习到的任务反馈来降低评估成本。我们提出了LESS(轻量级进化超网络搜索),一种数据驱动的方法,将简短的公平硬路径预热与单一CMA-ES分布下的离散搜索相结合。每个候选方案在六次候选条件超网络更新后,以其解码的硬基因型进行评估。在NAS-Bench-201上,LESS在409.1秒内达到了93.189±0.467%的CIFAR-10测试准确率,与FairNAS相差0.04个百分点,而搜索时间约为其源报告时间的1/24。匹配对照表明,校准将选中的验证准确率提高了0.577个百分点,而最佳访问准确率仅变化了0.054个百分点,表明其主要效果是减少选择遗憾。冻结配置无需调优即可迁移到CIFAR-100和ImageNet16-120,准确率分别为69.615±1.139%和43.720±1.697%。无需调优地应用于更大的DARTS搜索空间,LESS在CIFAR-10上达到96.95±0.14%,在CIFAR-100上达到82.43±0.80%,每次搜索在单个GPU上约43.5分钟完成。这些结果共同表明,简短、平衡、数据依赖的更新能够在几分钟内实现跨数据集和搜索空间的具有竞争力的神经架构搜索。
英文摘要
Low-cost NAS must both explore high-performing architectures and identify them reliably, yet reducing evaluation cost often weakens the fidelity of candidate comparisons. Training-free methods reduce evaluation cost by replacing learned task feedback with proxy signals measured at initialization. We introduce LESS (Lightweight Evolutionary Supernet Search), a data-driven method that combines a brief fair hard-path warm-up with discrete search under a single CMA-ES distribution. Each proposal is evaluated as its decoded hard genotype after six candidate-conditioned supernet updates. On NAS-Bench-201, LESS achieves \(93.189\pm0.467\%\) CIFAR-10 test accuracy in 409.1 seconds, coming within 0.04 percentage points of FairNAS using approximately \(1/24\) of its source-reported search time. Matched controls show that calibration improves selected validation accuracy by \(0.577\) percentage points while changing best-visited accuracy by only \(0.054\) points, indicating that its primary effect is to reduce selection regret. The frozen configuration transfers without tuning to CIFAR-100 and ImageNet16-120 with \(69.615\pm1.139\%\) and \(43.720\pm1.697\%\) accuracy. Applied without tuning to the larger DARTS space, LESS achieves \(96.95\pm0.14\%\) on CIFAR-10 and \(82.43\pm0.80\%\) on CIFAR-100, with each search completing in approximately 43.5 minutes on a single GPU. Together, these results show that short, balanced, data-dependent updates enable competitive neural architecture search across datasets and search spaces within minutes.
Comments33 pages, 4 figures. Code: https://github.com/AviralGandhi/LESS