arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Harness-Zero:通过智能体即框架(Agent-as-Harness)进行框架蒸馏

Harness-Zero: Harness Distillation via Agent-as-Harness

Haoran Ye, Yuxing Lu, Haonan Dong, Zhaochen Su, Guojie Song

arXiv 2609.24974首次发表:更新:

AI 中文总结

提出Harness-Zero,通过智能体即框架将优化框架的行为蒸馏进模型权重,部署时移除专用框架,将任务成功率从23.3%提升至44.3%,并恢复82.3%的框架诱导行为。

AI 中文摘要

智能体框架(agent harnesses)是协调模型与环境交互的外部系统,能够显著提升智能体性能,但其增益在部署时仍依赖于该框架。由于最佳框架因领域、实例和模型而异,通用智能体要么只能接受次优的共享框架,要么在日益增多的专用框架之间进行路由选择。因此,我们研究智能体框架蒸馏(agent harness distillation):将领域或实例优化的框架用作训练时指导,并将其诱导的行为迁移到模型权重中,从而使其增益在单一固定的目标框架下得以保留。挑战在于两种框架在动作空间和可用信息上存在差异,因此来自优化框架的指导不能直接作为目标框架的监督信号。我们提出了Harness-Zero,通过智能体即框架(agent-as-harness)实现框架蒸馏。在优化框架的指导下,一个“驾驭智能体”(harnessing agent)在目标框架的动作空间内、在执行前纠正学生响应,将框架指导转化为训练示范。对由此产生的轨迹进行微调,可将框架诱导的行为内化到模型中,从而在部署时可移除专用框架。我们在知识工作、工具使用和科学领域进行的实验表明:(1)对于使用相同进化框架的前沿大语言模型,智能体即框架优于代码即框架(code-as-harness)。(2)在部署时移除专用框架的情况下,Harness-Zero将基础模型的宏平均任务成功率从23.3%提升至44.3%,甚至超过了该框架仍保留时的41.7%。(3)Harness-Zero恢复了基础模型中缺失的框架诱导行为,在三个领域的28种模式中平均恢复率达82.3%。

英文摘要

Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness. The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one. We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness's action space, turning harness guidance into training demonstrations. Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment. Our experiments spanning knowledge work, tool use, and science domains show that: (1) For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness. (2) With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached. (3) Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑