EvoAlloc:一种用于高效程序演化的自演化资源分配智能体
EvoAlloc: A Self-Evolving Resource Allocation Agent for Efficient Program Evolution
浏览论文内容
中文总结 AI 辅助
针对现有LLM程序演化资源分配策略固定导致资源浪费的问题,提出自演化资源分配智能体EvoAlloc,通过整合搜索经验与反事实探索优化分配策略,在基准测试中大幅减少评估成本并提升性能。
中文摘要 AI 辅助
基于大语言模型(LLM)的程序演化依赖评估反馈来指导对高性能程序的迭代搜索,但评估通常计算成本高昂,因此必须将有限资源分配给能最有效推进搜索的候选程序。现有基于LLM的方法在整个搜索过程中通常采用固定分配策略,可能会在低价值候选程序上浪费资源,同时忽略有前景的候选程序。我们提出EvoAlloc,这是一种自演化资源分配智能体,它从搜索经验中学习,以修改其在候选程序间分配计算资源的策略。EvoAlloc会定期将先前的搜索和分配结果整合为可复用的经验,为后续的策略修订提供依据;它还采用反事实探索机制,偶尔评估分配器拒绝资源的候选程序,以揭示其结果,丰富未来策略更新的经验。在编码和智能体 harness 优化基准测试中,EvoAlloc达到基线性能所需的完整评估次数减少了59%-82%,总LLM token数减少了61%-89%;此外,在相同的完整评估预算下,EvoAlloc的最终性能提高了8.7%-12.0%。
英文摘要
LLM-based program evolution relies on evaluation feedback to guide the iterative search for high-performing programs. However, evaluation is often computationally expensive, making it essential to allocate limited resources to candidates that can most effectively advance the search. Existing LLM-based methods typically rely on fixed allocation strategies throughout the search, potentially wasting resources on low-value candidates while overlooking promising ones. We propose EvoAlloc, a self-evolving resource-allocation agent that learns from search experience to revise its strategy for allocating computational resources across candidates. EvoAlloc periodically consolidates prior search and allocation outcomes into reusable experience, which informs subsequent strategy revisions. It further uses a counterfactual exploration mechanism to occasionally evaluate candidates denied resources by the allocator, revealing their outcomes to enrich its experience for future strategy updates. Across coding and agent-harness optimization benchmarks, EvoAlloc requires 59-82% fewer full evaluations and 61-89% fewer total LLM tokens to reach baseline-level performance. Moreover, under the same full-evaluation budget, EvoAlloc achieves 8.7-12.0% higher final performance.
发表机构
- King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
- Sakana AI
机构由 AI 辅助整理,请以论文原文为准。