发表机构
National University of Singapore; University of Washington; Princeton University; Stanford University(新加坡国立大学; 华盛顿大学; 普林斯顿大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FrugalEvo提出成本感知的LLM引导程序演化框架,通过高低成本模型协作和缓存优化,在固定预算内最大化收益,在多个任务上以极低成本超越现有基线。
AI 中文摘要
LLM引导的演化方法,如AlphaEvolve,已成为解决具有挑战性的计算优化问题(如圆形填充)的强大方法。然而,先前的工作通常优化固定迭代次数内的性能提升。我们认为,实际优化应最大化每单位成本的收益。为此,我们提出了FrugalEvo,一个成本感知的演化框架,其中更强、成本更高的LLM探索解决方案策略,而更便宜的LLM实施这些策略并迭代改进生成的代码。我们还设计了一个缓存高效的演化过程,其中我们的框架和提示在不同演化步骤之间最大化前缀共享,以提高缓存复用。为了在固定成本预算内衡量解决方案质量,我们引入了预算感知曲线下面积(BA-AUC),定义为在累积LLM成本下,迄今为止最佳评估分数曲线下的面积,直至预算。在10个数学和系统优化任务中,FrugalEvo在最终解决方案质量上匹配或超越了最先进的基线,包括OpenEvolve、ShinkaEvolve、AdaEvolve和EvoX,并在9个任务上实现了更高的BA-AUC。它还在来自ALE-Bench-Lite的10个算法优化任务上实现了比这些基线更高的平均性能。值得注意的是,在圆形填充问题上,FrugalEvo使用GPT-5.6 Terra和Luna仅花费1.68美元,使用GLM-5.3及其Flash变体仅花费0.55美元,就实现了新的最先进性能,匹配或超越了所有基线,包括多智能体方法如CORAL和SwarmResearch,这些方法平均花费约50美元。
英文摘要
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the resulting code. We also design a cache-efficient evolution process, where our harness and prompts maximize the sharing of prefixes across different evolution steps, to improve cache reuse. To measure solution quality throughout a fixed cost budget, we introduce Budget-Aware Area Under the Curve (BA-AUC), defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget. Across 10 mathematical and systems optimization tasks, FrugalEvo matches or surpasses state-of-the-art baselines, including OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX, in final solution quality and achieves higher BA-AUC on 9 tasks. It also achieves higher average performance than these baselines on 10 algorithmic optimization tasks from ALE-Bench-Lite. Notably, on circle packing, FrugalEvo achieves new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD and with GLM-5.3 and its Flash variant for only 0.55 USD, matching or surpassing all baselines, including multi-agent methods such as CORAL and SwarmResearch, which cost approximately 50 USD on average.
Comments17 pages, 4 figures