发表机构
Cleer Science; The University of Edinburgh; Tsinghua University(克利尔科学公司; 爱丁堡大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对智能体自进化技能的缺陷,提出reSolve框架,通过求解与复现协议、代理验证器增强奖励及束搜索,使廉价模型技能性能超人工策划基线,且在自然科学任务上表现突出。
AI 中文摘要
智能体技能是智能体在部署时参考的可移植指令包和资源,目前自进化技能存在两个问题:一是从零进化的技能性能逊于人工策划的技能,且在弱模型上使用无技能的效果更差;二是进化阶段记录的一条幸运轨迹,新的随机智能体在部署时往往无法复现。我们提出reSolve,这是一种每任务、带专家回路的框架,包含三个组件:它将交互式求解与独立可交付成果解耦,该成果可在新容器中独立重新执行,我们将此协议称为求解与复现;它用无法访问隐藏测试或参考答案的代理验证器增强稀疏奖励信号;随后在解决方案构建图上运行验证器引导的束搜索。在固定框架内,廉价模型自进化的技能达到74.9%的3项均值,较60.1%的人工策划基线高出14.8个百分点,超过最强的官方策划技能结果(67.3%,GPT-5.5/OpenHands)。我们还报告了观察到的失败案例和领域级结果,包括在14项自然科学任务上的性能,以阐明该方法何时有效、何时无效。
英文摘要
Agent skills are portable packages of instructions and resources an agent consults at deployment. Self-evolving them fails in two ways today. First, skills evolved from scratch underperform human-curated ones and, on a weak model, using no skill at all. Second, an evolution-time pass records one lucky trajectory that a fresh stochastic agent often fails to reproduce at deployment. We present reSolve, a per-task, oracle-in-the-loop framework built on three components. It decouples interactive solving from a self-contained deliverable that is independently re-executed in a fresh container, a protocol we call solve-and-reproduce. It enhances the sparse reward signal with a surrogate verifier that cannot access hidden tests or reference answers. It then runs verifier-guided beam search over a solution-construction graph. Within a fixed harness, a cheap model self-evolves skills that reach $74.9\%$ mean-of-3, $+14.8$ points over the $60.1\%$ human-curated baseline, exceeding the strongest official curated-skill result ($67.3\%$, GPT-5.5/OpenHands). We also report observed failure cases and domain-level results, including performance on the 14 Natural Science tasks, to clarify when the approach does and does not help.