发表机构
Shanghai Jiao Tong University; University of California, Berkeley(上海交通大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 harness 退火训练(HAT),通过课程式弱化外部控制,使语言智能体在保留任务性能的同时内化控制决策,实验表明可减少运行时干预。
AI 中文摘要
语言智能体依赖外部 harness(控制框架)来跟踪状态、组织工作流程并验证答案。除了提供工具和信息外,这些 harness 还提供关于调查什么、是否修订以及何时停止的控制决策。在成功的 harness 支持轨迹上进行训练可以提高任务性能,同时将这些决策留给运行时干预。我们询问 harness 支持的经验是否也能教会模型做出这些决策,从而允许控制分工随着模型的学习而改变。我们将这一目标称为 harness 内化:学习承担指定的控制责任,同时在相应支持被撤销后保持任务性能。我们引入了 HARNESS ANNEALING TRAINING(HAT,即 harness 退火训练),它将显式控制监督与在逐渐弱化的 harness 下收集的教师轨迹上的课程相结合。在 SWE-QA 和 SWE-QA-Pro 上使用 9B 和 35B 模型进行的实验,在四种部署 harness 下评估了每个检查点。选定的退火检查点仅使用工具操作时,其得分接近各自在完整 harness 下部署的起始检查点的得分。收益随模型规模和部署配置而变化,进一步退火并不统一地提高性能。这些发现表明,harness 支持的经验可以帮助减少训练后的智能体所需的运行时控制。
英文摘要
Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories can improve task performance while leaving these decisions dependent on runtime intervention. We ask whether harness-supported experience can also teach the model to make these decisions, allowing the division of control to change as the model learns. We call this objective harness internalization: learning to assume specified control responsibilities while retaining task performance after the corresponding support is withdrawn. We introduce HARNESS ANNEALING TRAINING (HAT), which combines explicit control supervision with a curriculum over teacher trajectories collected under progressively weaker harnesses. Experiments with 9B and 35B models on SWE-QA and SWE-QA-Pro evaluate every checkpoint under four deployment harnesses. Selected annealed checkpoints operating with tools alone achieve scores close to those of their respective starting checkpoints deployed with the full harness. The benefits vary with model scale and deployment configuration, and further annealing does not uniformly improve performance. These findings suggest that harness-supported experience can help reduce the runtime control required by a trained agent.