发表机构
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Shenzhen University of Advanced Technology; University of Macau(中国科学院深圳先进技术研究院; 中国科学院大学; 深圳理工大学; 澳门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Growing Harness通过失败引导训练将重复控制转化为可复用代码,减少LLM调用并提升成功率,实现可复用的专家智能体。
AI 中文摘要
大型语言模型(LLM)智能体经常处理一系列相关任务,但标准工具架会在每个任务的上下文中反复要求模型重建相同的控制决策。我们研究任务反馈是否可以将重复性控制转化为可复用的可执行代码,同时将LLM调用保留给特定于任务的语义推理。我们引入了Growing Harness,一种失败引导的训练范式,它从无策略脚手架中学习智能体工具架本身,该脚手架暴露了固定的模型和工具接口,但不编码任何任务求解控制器。函数级执行轨迹将每个失败定位到有界的代码表面,优化器联合修复一组失败,成功优先的保留门控会回滚损害先前能力的修复序列。被接受的编辑累积在一个共享工具架中,使其控制结构能够从任务反馈中涌现。在BrowseComp-Plus和WebArena-Verified上,使用三种部署模型(参数从4B到120B),Growing Harness在六个基准-模型设置中的五个中取得了最高的平均成功率,并在第六个设置中落后最佳平均值0.7个百分点。相对于工具调用智能体,它减少了76.0-91.8%的LLM调用和74.4-98.6%的部署智能体推理成本。在WebArena-Verified上,其成功率在不同模型规模下保持在44.7-45.3%,而工具调用智能体在4B模型上降至6.7%。消融研究表明,轨迹局部编辑、联合修复和基于门控的回滚各自提升了最终成功率。这些结果表明,持续的程序增长可以将重复性控制从模型上下文中移出,转移到低成本代码中,产生可复用的专家智能体,这些智能体在使用较小部署模型时仍然有效。
英文摘要
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.
Comments16 pages, 6 figures