发表机构
University of Würzburg; Computer Vision Lab, CAIDAS & IFI(维尔茨堡大学; 计算机视觉实验室,CAIDAS与IFI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究用大语言模型扩展闭环特征通道配置,将搜索设置扩展到每个微调周期250个候选网络,分析大量候选网络,发现准确率呈正线性趋势,参数效率提高,还揭示了架构规律,表明通道搜索信号可转移。
AI 中文摘要
基于闭环大语言模型的通道配置搜索的初步成果表明,可以通过可执行代码生成和准确性反馈直接优化神经网络宽度。然而,这些结果来自相对稀疏的有效评估集,尚不清楚观察到的优化行为是否能转移到更密集的采样机制,以及在评估更多生成的网络时是否会出现其他架构规律。为了测试这一点,将相同的搜索设置扩展到每个微调周期250个候选网络。分析涵盖了来自8个完整周期的2000个生成候选网络,经过任务和元数据过滤后得到462个经过验证的CIFAR-100评估结果。每个周期的平均准确率呈现正线性趋势,斜率为9.87e-4(p=0.043),而高性能前沿的提升更为显著:最佳观察到的准确率从0.3144提高到0.3676,前5和前10周期级别的平均值均呈现正趋势。扩展运行还显示出参数效率的提高。最佳模型在1180万个参数时达到0.3676,而早期高性能模型在1.665亿个参数时为0.3144。除了准确率,更大的样本还揭示了从稀疏观察中难以评估的架构规律。41.8%的经过验证的候选网络出现非2的幂次方通道宽度,最强的模型共享结构化通道分配模式,其特点是早期宽度适中,中间或后期块扩展。这些发现表明,在初步研究中观察到的通道搜索信号可以转移。
英文摘要
Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimized directly through executable code generation and accuracy feedback. However, those results were obtained from a relatively sparse set of valid evaluations, leaving open whether the observed optimization behavior transfers to a denser sampling regime and whether additional architectural regularities emerge when more generated networks are evaluated. To test this, the same search setting is scaled to 250 candidate networks per fine-tuning cycle. The analysis covers 2000 generated candidates from 8 complete cycles, yielding 462 verified CIFAR-100 evaluations after task and metadata filtering. Per-cycle mean accuracy exhibits a positive linear trend with slope 9.87e-4 (p=0.043), while the high-performing frontier improves more strongly: the best observed accuracy increases from 0.3144 to 0.3676, and both the top-5 and top-10 cycle-level means exhibit positive trends. The scaled run also reveals improved parameter efficiency. The best model reaches 0.3676 with 11.8M parameters, compared with an early high-performing model at 0.3144 with 166.5M parameters. Beyond accuracy, the larger sample exposes architectural regularities that were difficult to assess from sparse observations. Non-power-of-two channel widths occur in 41.8% of verified candidates, and the strongest models share structured channel-allocation patterns characterized by moderate early widths and expanded middle or later blocks. These findings indicate that the channel-search signal observed in the initial study transfers
Comments15 pages, 8 figures