发表机构
Shanghai Jiao Tong University; Xiaohongshu Inc.; Southeast University(上海交通大学; 小红书公司; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Harness-R1是首个基于智能体失败轨迹学习编辑可执行运行时管控程序的方法,经WebShop等基准测试,可提升Qwen3.5-9B智能体成功率,为管控程序与智能体协同进化提供了方向。
AI 中文摘要
基于大语言模型构建的智能体在部署过程中会不断积累交互轨迹,但其行为通常保持固定。除了更新模型权重外,这些轨迹还可以改进用于构建上下文、协调工具、验证动作及恢复执行的智能体管控程序。我们提出Harness-R1,据我们所知,这是首个能对现有可执行运行时进行基于失败的全生命周期编辑的学习方法。该方法通过在线强化学习对专用的管控程序工程师进行后训练,使其编辑操作以实际任务成功为优化目标,而非由固定编辑器提出。另一个9B规模的工程师将目标智能体的批量失败转换为经过验证的可执行补丁;对冻结的目标智能体进行同批次的重新运行以提供结果奖励,因此训练仅更新工程师。冷启动阶段的监督微调初始化该编辑策略,随后使用组相对策略优化进行在线训练。在WebShop、ALFWorld和DBBench三个基准测试中,Harness-R1将普通Qwen3.5-9B的成功率从44.3%提升至53.6%,增幅9.3个百分点;在对目标智能体进行直接微调后,针对特定目标的工程师进一步将平均成功率从59.2%提升至64.2%,增幅5.0个百分点;由于这些增益在目标智能体微调前后均成立,Harness-R1为管控程序工程师与目标智能体的协同进化指明了方向。
英文摘要
Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weights, these trajectories can improve the agent harness that constructs context, mediates tools, validates actions, and recovers execution. We introduce Harness-R1, the first method, to our knowledge, that makes failure-conditioned, lifecycle-wide editing of an existing executable runtime a learned capability. It post-trains a dedicated harness engineer with online reinforcement learning so that its edits are optimized for the realized task success they produce, rather than proposed by a fixed editor. A separate 9B engineer converts batches of target-agent failures into validated executable patches; fresh same-batch reruns of the frozen target provide outcome rewards, so training updates only the engineer. Cold-start supervised fine-tuning initializes this editing policy, which is then trained online with group-relative policy optimization. Across WebShop, ALFWorld, and DBBench, Harness-R1 raises vanilla Qwen3.5-9B success from 44.3% to 53.6% (+9.3 percentage points). After direct target-agent fine-tuning, a target-specific engineer raises the average further from 59.2% to 64.2% (+5.0 points); because these gains hold both before and after fine-tuning the target, Harness-R1 points toward co-evolving the harness engineer and the target agent.