arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自博弈与技能演化:能提出、解决并记忆的自演化搜索智能体

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, Changwei Wang

arXiv 2607.29468首次发表:更新:

AI 中文总结

本研究提出SESA智能体,通过任务生成与技能记忆的双向循环协同演化,在7个问答基准上较SSP、SkillRL等基线实现准确率提升,且支持无记忆部署与可选推理检索。

AI 中文摘要

自博弈智能体可在无目标基准问题的情况下生成训练问题,但其课程缺乏持久状态:失败会影响梯度,但不会明确塑造后续实践。外部技能记忆保留过程经验,但通常从固定任务分布中学习。我们提出SESA(自演化技能增强智能体),它将过程记忆转化为工具增强搜索自博弈的演化状态:挑战者提出问题,单独参数化的求解器仅检索技能;将有价值的失败提炼为可复用技能并写回记忆,更新后的记忆会改变求解器行为与成功率,进而改变挑战者的奖励及未来问题分布,生成的新失败会再次改写记忆。这种双向循环使任务生成与技能记忆协同演化。由于检索到的技能会影响在线策略训练轨迹,其益处既可进入模型参数,也可保留在外部库中,支持无记忆部署与可选推理时检索。在7个开放域多跳问答基准上,SESA在多个主干模型上的平均准确率较SSP提升1.2-3.2个百分点,在统一评估协议下较技能增强的SkillRL基线提升0.9个百分点;在Qwen3模型上,SESA-Off较SSP仍保留1.8-2.2个百分点的提升,最终技能库额外贡献0.5-1.0个百分点。这些结果表明,演化技能记忆不仅是推理时的插件,还会改变策略学习与未来训练分布,同时作为可选外部记忆保留价值。代码可在该https URL获取。

英文摘要

Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice. External skill memories preserve procedural experience but are typically learned from fixed task distributions. We introduce \textbf{SESA} (Self-Evolving Skill-Augmented Agent), which makes procedural memory an evolving state of tool-augmented search self-play. A challenger poses problems, while a separately parameterized solver alone retrieves skills. Informative failures are distilled into reusable skills and written back to memory. The updated memory changes solver behavior and success, which changes the challenger's reward and the distribution of future problems; the resulting frontier produces new failures that rewrite memory. This bidirectional loop makes task generation and skill memory co-evolve. Because retrieved skills shape on-policy training trajectories, their benefits can enter the model parameters as well as remain in the external bank, enabling memory-free deployment and optional inference-time retrieval. Across seven open-domain and multi-hop question-answering benchmarks, SESA improves average accuracy over SSP by 1.2--3.2 points across multiple backbones and surpasses the skill-augmented SkillRL baseline by 0.9 points under a unified evaluation protocol. On Qwen3 models, SESA-Off retains 1.8--2.2 points of improvement over SSP, while the final skill bank adds a further 0.5--1.0 points. These results show that evolving skill memory is not merely an inference-time plug-in: it changes policy learning and the future training distribution while retaining value as optional external memory. Our code is available at https://github.com/Zenghuang-Fu/SESA-Self-Evolving-Search-Agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑