发表机构
Czech Technical University(捷克技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出在变分自编码器学习的潜在空间中应用贝叶斯优化进行零样本神经架构搜索,通过代理标量化组合多个零样本代理,在10,000次迭代内找到在图像分类、目标检测和语义分割上达到最先进结果的架构。
AI 中文摘要
零样本神经架构搜索消除了传统神经架构搜索的昂贵成本,但其搜索过程通常基于进化算法(EA);由于缺乏对目标的显式建模,它常常通过变异进行近乎随机的搜索。贝叶斯优化提供了一种原则性的替代方案,通过建模目标并跨迭代聚合信息,但在现代神经架构搜索的高维、离散、图结构空间中扩展性差,限制其仅适用于小型网络。在本文中,我们通过变分自编码器学习一个潜在空间,该自编码器训练用于重构一种新颖的架构前缀编码,从而将贝叶斯优化引入大规模架构的零样本神经架构搜索,并提出一种代理标量化方法,将多个零样本代理组合成单一的贝叶斯优化目标。在仅经过所提搜索算法的10,000次迭代(在单个GPU上耗时8小时)后,我们的方法在给定的模型参数数量约束下,找到了一个在三个独立任务——图像分类、目标检测和语义分割——上达到最先进结果的网络架构。
英文摘要
Zero-shot Neural Architecture Search removes the prohibitive cost of traditional NAS, but its search process is typically based on the evolutionary algorithm (EA); lacking an explicit model of the objective, it often resorts to a near-random search through mutation. Bayesian Optimization offers a principled alternative by modeling the objective and aggregating information across iterations, but scales poorly to the high-dimensional, discrete, graph-structured spaces of modern NAS, restricting its use to only small networks. In this paper, we bring Bayesian Optimization to zero-shot NAS for large-scale architectures by learning a latent space via a Variational Autoencoder trained to reconstruct a novel prefix encoding of architectures and propose a proxy scalarization that combines several zero-shot proxies into a single Bayesian Optimization objective. After only 10,000 iterations of the proposed search algorithm (8 hours on a single GPU), our method found a network architecture which under the given model parameter count constraints achieves state-of-the-art results on three separate tasks -- image classification, object detection and semantic segmentation.