AdaEva:利用自适应部分评估加速LLM驱动的算法设计
AdaEva: Accelerating LLM-Driven Algorithm Design with Adaptive Partial Evaluation
浏览论文内容
中文总结 AI 辅助
AdaEva提出自适应部分评估框架,通过逐步淘汰不具前景的候选算法,在保持LLM4AD流程不变的情况下,显著提升评估效率与搜索性能,并改善泛化能力。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被用于自动化算法设计。然而,评估所生成算法的计算成本可能过高。我们考虑常见的LLM驱动的自动化算法设计(LLM4AD)设置,其中候选算法通过在共享的训练实例集上聚合其性能来进行评估。这种实例级结构引发了一个自然的问题:在决定候选算法是否仍然具有竞争力之前,是否必须对每个候选算法在整个实例集上进行评估?受算法配置的启发,我们引入了AdaEva,一个即插即用的自适应部分评估框架,它逐步在相同实例池的更大子集上评估候选算法,并随着证据的积累淘汰不具前景的候选算法。重要的是,AdaEva保持底层LLM4AD过程和每个实例的评估器不变,并且不需要关于实例难度的先验知识。我们使用连续减半(AdaEva-S)和统计竞赛(AdaEva-R)来实例化这一想法,并在三个代表性的LLM4AD框架、多个LLM骨干网络以及涵盖组合和连续黑盒优化的优化领域中对这两种机制进行了评估。在匹配的评估预算下,AdaEva比固定的部分评估策略更可靠地在候选算法之间平衡评估工作量,从而在评估的设置中产生强大的搜索效率和任意时间性能,以及改进的留出泛化能力。
英文摘要
Large Language Models (LLMs) are increasingly used for automated algorithm design. However the computational cost of evaluating the generated algorithms can be excessive. We consider the common LLM-driven automated algorithm design (LLM4AD) setting in which a candidate algorithm is evaluated by aggregating its performance over a shared set of training instances. This instance-wise structure raises a natural question: must every candidate be evaluated on the entire instance set before deciding whether it remains competitive? Taking inspiration from algorithm configuration, we introduce AdaEva, a drop-in adaptive partial-evaluation framework that progressively evaluates candidates on larger subsets of the same instance pool and eliminates unpromising candidates as evidence accumulates. Importantly, AdaEva leaves the underlying LLM4AD procedure and per-instance evaluator unchanged and requires no prior knowledge about instance difficulty. We instantiate this idea using successive halving (AdaEva-S) and statistical racing (AdaEva-R), and evaluate both mechanisms across three representative LLM4AD frameworks, multiple LLM backbones, and optimization domains spanning combinatorial and continuous black-box optimization. Under matched evaluation budgets, AdaEva more reliably balances evaluation effort across candidates than fixed partial-evaluation strategies, yielding strong search efficiency and anytime performance together with improved held-out generalization across the evaluated settings.
发表机构
- University of St Andrews(圣安德鲁斯大学)
- Sorbonne Université(索邦大学)
- CNRS(法国国家科学研究中心)
- LIP6(计算机科学实验室(LIP6))
- University of Zurich(苏黎世大学)
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。