发表机构
The University of Alabama(阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EvE是一种结合差分进化与针对性Adam回退的优化器,在预算受限的搜索场景中比Adam快数倍,且排名一致性相当,但最终质量略有损失。
AI 中文摘要
Adam及其变体主导着神经网络训练,但单次运行只有在大部分预算消耗后才能揭示配置是否有效,这不适于超参数或架构搜索,因为在这些场景中配置必须被廉价地排名并提前剪枝。我们提出EvE(进化探索器),一种稳态的、种群大小为四的差分进化(DE)优化器,并带有针对性的Adam回退机制:每次迭代通过DE提出一个候选,仅当DE步骤未能改进当前最优解时,才运行一小段梯度下降。选择是贪婪的,因此在确定性目标上,迄今为止最优值可证明是单调非增的,并且由于梯度仅用作针对性救援,每次迭代的成本保持在单个Adam步骤的常数因子内,与维度无关。在固定的、评估成本匹配的预算下,EvE在七个可扩展基准(最高达一百万个变量)的70个(问题,维度)单元中,有76%的单元胜过或持平Adam。在三个真实神经网络任务(MNIST上的MLP,以及两个数据集上1.5B参数语言模型的LoRA微调)中,EvE在相同的计费预算下完成速度快1.7-3.9倍,但最终质量略有损失(MNIST上约一个准确率点,两个微调任务上相对测试损失高9-11%;在GSM8K上,Adam准确率高约5个点,且微调使两者的准确率均低于基础模型)。在UCI Adult上的连续减半中,EvE完成超参数和架构搜索快3.1-3.5倍,配置排名与Adam自身跨种子的排名一致性相当(Kendall's tau 0.66-0.69)。EvE并非完全替代Adam作为最终阶段训练器,而是一种快速、梯度感知的代理,适用于更高层次的、搜索密集且预算受限的场景。
英文摘要
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and since gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven scalable benchmarks up to one million variables. On three real neural-network tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE finishes the same charged budget 1.7-3.9x faster, at a modest cost in final quality (about one accuracy point on MNIST, 9-11% higher relative test loss on the two fine-tuning tasks; on GSM8K, Adam is about 5 accuracy points more accurate, and fine-tuning lowers accuracy below the base model for both). Inside successive halving on UCI Adult, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, ranking configurations about as consistently with Adam as Adam does with itself across seeds (Kendall's tau 0.66-0.69). EvE is not a total replacement for Adam as a final-stage trainer, but a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up.