arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05651cs.CLcs.AIcs.NE

中继而非路由:面向成本高效的大语言模型驱动演化的自适应种群交接

Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution

Sichun Luo, Yi Huang, Guanzhi Deng, Haibo Wang, Haochen Luo, Lei Li, Zefa Hu, Junlan Feng, Qi Liu

AI总结:

针对LLM驱动演化全程用强模型成本高的问题,提出无需训练的自适应种群交接框架,通过廉价模型探索、中继增益触发交接,在12个测试设置中11个获最高平均分数,性能优于基线。

AI中文摘要:

大语言模型(LLM)驱动的演化在程序搜索和算法发现领域已展现出应用前景,但在漫长的演化过程中全程依赖强模型的成本较高。一个自然的替代方案是在固定推理预算下将廉价模型与强模型结合使用。然而,现有方法通常在单个查询或变异步骤层面分配模型,却忽略了演化搜索具有「有状态」的特性:每个生成的候选都会改变后续变异所依赖的种群。我们对LLM驱动的演化轨迹进行了实证分析,发现搜索进展具有明显的前置性,早期轨迹性能具有参考价值但存在噪声,且廉价模型能以更低成本恢复强模型实现的大部分早期进展。基于这些发现,我们提出了\textbf{\textbf{\textit{模型名}}},这是一个无需训练的框架,通过自适应「种群交接」将预算分配从单个调用转向演化中的种群。廉价模型在由多臂老虎机调度器分配的短时间块中探索多条轨迹;中继增益被定义为为交接构建的紧凑、质量多样的候选库的边际改进,用作调度器的奖励并决定交接时机;经筛选的候选将初始化一个共享的强模型种群以进行优化。在四个基准测试和三种预算设置中,\textbf{\textbf{\textit{模型名}}}在12个设置中的11个里取得了最高平均分数,优于具有竞争力的基线方法。我们的结果表明,在有状态搜索中,预算分配应围绕种群而非单个调用组织。

英文摘要:

Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget. However, existing approaches typically allocate models at the level of individual queries or mutation steps, overlooking that evolutionary search is \textit{stateful}: each generated candidate changes the population from which subsequent mutations are produced. We empirically analyze LLM-driven evolutionary trajectories and find that search progress is strongly front-loaded, early trajectory performance is informative but noisy, and cheap models recover much of the early progress achieved by strong models at lower cost. Motivated by these findings, we propose \textbf{\model}, a training-free framework that shifts budget allocation from individual calls to evolving populations through adaptive \textit{population handoff}. A cheap model explores multiple trajectories in short blocks allocated by a bandit scheduler. Relay Gain, defined as the marginal improvement of a compact, quality-diverse candidate bank constructed for handoff, serves as the scheduler reward and determines when to hand off. The curated candidates initialize a shared strong model population for refinement. Across four benchmarks and three budgets, \model achieves the highest mean score in 11 of 12 settings, outperforming competitive baselines. Our results suggest that in stateful search, budget allocation should be organized around the population, not the individual call.

↑