arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11826cs.LGcs.AIcs.NE

用于节俭神经架构搜索的Transformer引导的群体智能

Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search

Romain Amigon

AI总结:

本文提出基于Transformer与人工蜂群算法结合的节俭神经架构搜索框架,通过动态熵机制防过早收敛,解决“冷启动”问题。在CIFAR-10及信用卡欺诈检测任务中表现出色,能发现高效、参数少的模型,适合边缘部署。

AI中文摘要:

神经架构搜索(NAS)实现了深度学习模型设计的自动化,但传统上需要大量计算资源,通常以数千个GPU日来衡量。本文提出了一个节俭且带有记忆性的NAS框架,旨在使消费级硬件上的架构设计民主化。我们的方法将通过强化学习训练的自回归Transformer控制器的全局宏观搜索能力与人工蜂群(ABC)算法的局部微观利用相结合。为防止强化学习阶段的过早收敛,我们引入了动态熵机制,在检测到性能停滞时强制进行拓扑探索。在标准GPU(NVIDIA RTX 3060)上进行评估,我们的混合方法有效解决了元启发式算法中固有的“冷启动”问题。通过算法惩罚网络深度,我们的框架积极减轻模型膨胀:在CIFAR-10数据集上,它在3小时的搜索时间内发现了一个高效架构,仅约174,000个参数就达到了84.85%的准确率(远小于像ResNet-20这样的标准基线)。此外,我们通过将其应用于信用卡欺诈检测来展示框架的灵活性,在高度不平衡的表格数据上直接优化F1分数,使用约4,600个参数的紧凑网络达到了0.71的F1分数。这些结果表明,我们的方法可以产生适合边缘部署的定制、可访问且高度参数高效的深度学习模型。

英文摘要:

Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive computational resources, often measured in thousands of GPU-days. In this paper, we propose a frugal and memetic NAS framework designed to democratize architecture design on consumer-grade hardware. Our approach combines the global macro-search capabilities of an autoregressive Transformer controller, trained via Reinforcement Learning (RL), with the local micro-exploitation of an Artificial Bee Colony (ABC) algorithm. To prevent premature convergence during the RL phase, we introduce a dynamic entropy mechanism that forces topological exploration upon detection of performance stagnation. Evaluated on a standard GPU (NVIDIA RTX 3060), our hybrid method effectively resolves the "cold-start" problem inherent in metaheuristics. By algorithmically penalizing network depth, our framework actively mitigates model bloat: on the CIFAR-10 dataset, it discovers an efficient architecture reaching 84.85% accuracy with only $\sim$174,000 parameters (significantly smaller than standard baselines like ResNet-20) in 3 hours of search time. Furthermore, we demonstrate the framework's flexibility by applying it to credit card fraud detection, directly optimizing the F1-Score on highly imbalanced tabular data to reach a F1-Score of 0.71 with a compact network of $\sim$4,600 parameters. These results suggest that our approach can yield tailored, accessible, and highly parameter-efficient deep learning models suitable for edge deployment.

↑