arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37968cs.AIcs.CLcs.LG

SelfSearch:面向自我改进智能体的无奖励搜索

SelfSearch: Reward-Free Search for Self-Improving Agents

Jungwoo Yang, Injin Kong, Yohan Jo

首次发表
浏览论文内容

中文总结 AI 辅助

SelfSearch是一种无奖励搜索方法,利用智能体自我修改的历史记录来改进自身,在多个基准上提升成功率并降低成本,且搜索成本极低。

中文摘要 AI 辅助

大语言模型智能体在编码能力上的进步使其能够检查和修改自身的指令、工具及执行流程。现有方法利用这一能力,通过重复的下游评估来搜索改进后的智能体,这会产生高昂的成本,并将搜索与所评估的任务绑定。我们提出SelfSearch,一种无奖励的搜索流程,智能体利用先前自我改进过程的记录来修改自身。这些记录捕获了先前修改尝试中的推理、工具操作及结果,为改进任务解决能力和自我修改提供了具体经验。在搜索过程中不依赖下游奖励信号的情况下,SelfSearch在所有六个模型-基准组合中均提升了初始智能体的总体平均成功率,其中在Terminal-Bench 2.1上,单个智能体的成功率最高提升了11.2个百分点。在SWE-bench Multilingual上,一个智能体在初始与进化智能体均能解决的任务上,成功率提升了5.0个百分点,同时执行成本降低了38.5%。SelfSearch在较低搜索成本下达到了与评估引导搜索基线相当的任务成功率。仅花费4.03美元的搜索成本,它就在公开的九种工具包对比设置下,使用DeepSeek V4 Flash解决了Terminal-Bench 2.1中82.0%的任务,与得分最高的工具包Codex持平。这些结果表明,通过自我修改获得的经验能够提升智能体的下游能力和效率。

英文摘要

Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.

发表机构

  • Graduate School of Data Science, Seoul National University(首尔大学数据科学研究生院)

机构由 AI 辅助整理,请以论文原文为准。

↑