AI 中文总结
ε-MemEvo是面向LLM程序进化的自适应跨任务记忆迁移框架,在8个优化基准上以GPT-5为主干,较AdaEvolve实现AUCC与早期收敛提升,且计算开销极低。
AI 中文摘要
基于大语言模型(LLM)的程序进化系统如FunSearch和AlphaEvolve已展现出发现新型算法的强大能力,但这类系统通常会孤立优化每个任务,任务完成后便丢弃搜索经验。本文提出ε-MemEvo,一种用于LLM程序进化的跨任务知识迁移框架。ε-MemEvo将过往经验存储为与任务无关的策略记忆,即成功算法策略的紧凑自然语言摘要而非原始代码,从而实现跨不同应用程序编程接口(API)和评估器的任务迁移。为避免语义不匹配记忆带来的负迁移,ε-MemEvo采用自适应注入门,用于决定是否注入检索到的记忆及注入强度。我们在涵盖数学优化和系统工程的8个不同优化基准上评估ε-MemEvo,采用排除目标任务记忆条目的内容级留一法(Leave-One-Out)协议。以GPT-5为主干模型,ε-MemEvo在全部8个任务上的AUCC指标均优于AdaEvolve,平均相对提升达+8.7%,早期收敛速度平均提升+9.4%。 ablation实验显示,单纯的记忆注入可能导致灾难性失败,而自适应门控在全部5个 ablation任务中均保持安全。数据更新后的后验在观测状态中具有可解释性:在搜索优化阶段倾向于弃权(不执行),并在早期和后期平台期从弃权转向提示。这些性能提升带来的计算开销不足1%。
英文摘要
LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically optimize each task in isolation, discarding search experience after completion. We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledge transfer in LLM program evolution. $\varepsilon$-MemEvo stores prior experience as task-agnostic tactic memories: compact natural-language summaries of successful algorithmic strategies rather than raw code, enabling transfer across tasks with different APIs and evaluators. To avoid negative transfer from semantically mismatched memories, $\varepsilon$-MemEvo uses an adaptive injection gate that decides whether retrieved memories should be injected, and at what intensity. We evaluate $\varepsilon$-MemEvo on 8 diverse optimization benchmarks spanning mathematical optimization and systems engineering, using a content-level Leave-One-Out protocol that excludes target-task memory entries. On the primary GPT-5 backbone, $\varepsilon$-MemEvo improves AUCC over AdaEvolve on all 8 tasks, with a mean relative gain of +8.7%, and improves early-stage convergence by +9.4% on average. Ablations show that naive memory injection can fail catastrophically, while adaptive gating remains safe across all five ablation tasks. The data-updated posterior is interpretable in observed states: it favors skip during improving search and shifts from skip to hint across early and late plateaus. These gains incur less than 1% computational overhead.