arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

进化还是错觉?重新思考大语言模型进化搜索中的评估

Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search

Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay

arXiv 2609.19799首次发表:更新:

发表机构

IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文质疑LLM进化搜索中单一预算评估的可靠性,通过网格实验揭示最佳种子-迭代分配随策略、任务和预算变化,并提出前沿测量协议。

AI 中文摘要

大语言模型驱动的进化搜索通过启动种子并迭代每个种子来寻找程序。论文通常报告单一的预算设置,通常是在固定迭代次数下运行一次种子,并据此对方法进行排名。我们表明这并不足够。我们在五个优化任务上评估了三种进化搜索策略,这些任务是该领域论文常用来报告结果的。我们在种子和迭代的完整网格上进行了分析。我们的研究结果表明,在更多种子(宽度)和更多迭代(深度)之间分配固定预算的最佳方式会随着策略、任务和总预算的变化而变化。此外,我们观察到策略的排名也随预算变化。在一个任务上,在一种种子下表现最差的策略在四十种种子下表现最佳。在另一个任务上,最佳迭代次数远低于实践中常见的值,因此额外的深度浪费了预算,而这些预算如果用于更多种子则会转化为得分。我们提供了一种测量协议,报告种子乘以迭代的前沿,并给出了使用它的实用指南。

英文摘要

LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterations (depth) changes with the strategy, the task, and the total budget. Furthermore, we observe that the ranking of strategies also changes with the budget. On one task the strategy that looks worst at one seed is best at forty seeds. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score. We provide a measurement protocol that reports the seeds-by-iterations frontier and practical guidance for using it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑