arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从研究中学习:迈向终身智能体框架演化

Learning from Research: Toward Lifelong Agent Harness Evolution

Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang

arXiv 2609.40169首次发表:更新:

发表机构

University of California, Santa Barbara; Microsoft(加州大学圣塔芭芭拉分校; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语言智能体持续改进需求,提出ScholarEvolve框架,利用研究文献自动演化智能体框架,在AppWorld和Tau2-Bench上显著提升任务完成率。

AI 中文摘要

语言智能体被期望解决日益复杂的任务,这产生了对持续改进的不断增长的需求。一种有前景的方法是演化智能体框架(agent harness),即控制工具使用、记忆管理和任务执行的软件,同时保持底层语言模型固定不变。近期方法通过使用元编码智能体基于执行反馈修改框架来自动化这一过程。然而,依赖该智能体的现有知识和观察到的失败可能会限制探索,并使适应变得被动。受人类专家从研究文献中学习新解决方案的启发,我们引入了ScholarEvolve,一个自动利用最新研究来指导框架演化的框架。ScholarEvolve将框架演化方向组织为功能模块,并使用主题建模来识别每个模块的不同改进策略。它实现这些策略并评估其组合以提升任务性能。此外,该框架设计为随时间纳入新出版物,使研究进展能够驱动主动的终身演化。实验在AppWorld和Tau2-Bench上展示了改进。ScholarEvolve将Qwen3.5-27B在AppWorld Challenge上的任务目标完成率从49.6%提升至63.6%,并将GPT-5.4-mini在Tau2-Bench Telecom上的pass@1从72.7%提升至81.9%。

英文摘要

Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑