发表机构
University of California, Berkeley; Purdue University; Stanford University(加利福尼亚大学伯克利分校; 普渡大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对长 horizon LLM 智能体,提出通过区分过程与内容失败,选择工具链进化或权重训练提升性能,在 DeepPlanning、WebArena-Lite 基准验证了方法有效性。
AI 中文摘要
提升长 horizon 大语言模型(LLM)智能体的性能,可通过在冻结模型周围进化工具链(harness)或对其权重进行训练实现。本文先让自进化工具链提升系统性能,再将进化后的工具链与基础模型及训练后的权重交叉结合,以探究训练后的模型保留哪些增益、哪些仍需运行时工具链支持。研究表明,可通过智能体的失败构成确定合适的调控手段:按首次触发信号标记失败轨迹,可将过程失败(被阻止的调用、循环、耗尽的步骤预算)与内容失败(生成的计划质量差)区分开。工具链进化可修复过程失败,其注入的行为可被训练到权重中,而内容失败正是权重训练的目标。在 DeepPlanning 基准测试中,自进化工具链循环将 Qwen3.5-4B 的保留得分从 0.16 提升至 0.30,将 Qwen3.5-9B 的保留得分从 0.32 提升至 0.44;对于 4B 模型,保留交付率从 55% 升至 90%,内容失败则留待权重处理。在进化工具链轨迹上训练的 LoRA 适配器可内化该增益:在原始工具链下,两种规模的模型在保留任务上均提升 0.13;对于 4B 模型,其与工具链叠加后保留得分翻倍以上,对于 9B 模型,适配器单独即可匹配完整进化线,将内容失败从四分之一的轨迹降至二十分之一。在答案打乱轨迹上训练的安慰剂适配器性能低于基础模型。该循环可迁移至 WebArena-Lite 基准(117 个未见过任务提升 0.09),此处增益存在于模型所见内容中,适配器无法进一步提升。最终形成两次应用的诊断-干预规则:读取失败构成以选择工具链或权重,再读取接受的编辑内容以决定训练哪些增益。得分是针对新基准的四次 rollout 均值,除标注外均为同夜测试,涉及六个系列的八个模型及两个基准。
英文摘要
Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.
Comments21 pages, 7 figures, 15 tables