发表机构
Salesforce AI(Salesforce人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出在策略专家校正流水线,解决工具包进化与模型微调冲突,使较弱模型在七个企业任务上性能提升。
AI 中文摘要
智能体工具包(围绕模型的系统提示、工具集、执行钩子和上下文管理脚手架)是智能体任务成功的关键决定因素。自动化工具包进化能使较小的模型以前沿模型成本的一小部分在特定领域任务上表现良好。由于工具包和模型权重共同塑造行为,我们探讨如何结合工具包进化与轻量级微调。在七个企业智能体任务中,我们首先用较弱模型进化工具包,然后发现较强的专家通常能更有效地使用它,这表明专家监督可以弥合剩余的差距。然而,在进化后的工具包下,用专家的完整轨迹训练较弱模型会适得其反:在Qwen3-Coder和Gemma 4上,所有七个任务的性能均下降4到30个百分点,尽管同样的程序在未进化的工具包下是有益的。我们的分析表明,模仿转移了知识并增加了脚手架使用,但破坏了模型与工具包的契合度:较弱模型采用了专家的规划策略却没有执行能力,不再匹配围绕其原生规划风格进化的工具包。因此,我们开发了一个在策略专家校正流水线,由元级MLE智能体自动化,定位较弱模型自身轨迹中失败的一轮,并让专家仅重写该轮。这保留了模型的规划风格,并结合了工具包进化和模型适应的收益。我们的结果识别并解决了工具包与权重更新之间的冲突源,为特定领域企业任务的经济协同进化提供了一种保持兼容性的方案。
英文摘要
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harness and model weights shape behavior, we ask how harness evolution and lightweight fine-tuning should be combined. Across seven enterprise agent tasks, we first evolve a harness with the weaker model, then find that a stronger expert often uses it more effectively, suggesting expert supervision could close the remaining gap. However, training the weaker model on the expert's complete trajectories under the evolved harness backfires: performance regresses on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, even though the same procedure helps under the unevolved harness. Our analysis shows that imitation transfers knowledge and increases scaffold usage, but disrupts model-harness fit: the weaker model adopts the expert's planning strategy without the competence to execute it and no longer matches the harness evolved around its native planning style. We therefore develop an on-policy expert-correction pipeline, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn. This preserves the model's planning style and combines the gains of harness evolution and model adaptation. Our results identify and resolve a source of contention between harness and weight updates, yielding a compatibility-preserving recipe for economical co-evolution on domain-specific enterprise tasks.