arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编写、执行、优化:通过执行反馈强化学习从技能跟随者到技能优化器

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

Kang Peng, Zhiwei Zhang, Yichen Zhang, Zezhong Wang, Yiming Du, Geng Tu, Baojun Wang, Bin Liang, Ruifeng Xu, Kam-Fai Wong

arXiv 2608.17587首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; The Chinese University of Hong Kong; Huawei Technologies Co., Ltd.; Harbin Institute of Technology(哈尔滨工业大学(深圳); 香港中文大学; 华为技术有限公司; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出WER多阶段框架,通过冻结执行器、程序验证器打分及混合轨迹优化训练技能优化器,在BFCL v4和tau2-bench上显著提升智能体工具使用性能,4B优化器表现优于通用模型。

AI 中文摘要

专家编写的自然语言技能可提升使用工具的智能体性能,但智能体自行编写的技能表现比不使用技能时差8-11个百分点。这一差距表明,遵循程序指导与基于执行证据改进指导是两种不同的能力。推理时的循环可修复技能,但无法改进编写下一个技能的模型。本研究探讨如何将中间技能的执行经验组织为优化器的训练状态。我们提出WER(Write, Execute, and Refine)这一多阶段框架,在冻结执行器之外训练技能优化器:优化器提出技能,冻结智能体对每个技能重复执行,程序验证器对结果打分;分数提供相对信用并选择混合结果记录,这些记录中匹配的成功与失败轨迹构成下一阶段的优化状态,使优化器能从自身早期输出的后果中学习。在BFCL v4多轮任务和tau2-bench上,WER相比无技能基准分别将平均Pass@1提升7.80和3.85个百分点;在相同优化流程下,其性能比未进行优化器训练的同主干模型分别高出9.35和10.29个百分点。训练后的4B规模优化器在BFCL v4上达到76.63%的准确率,平均表现优于所有被评估的现成通用技能优化器模型。

英文摘要

Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improve the model that writes the next one. We study how to organize execution experience from intermediate skills into training states for an optimizer. We introduce WER (Write, Execute, and Refine), a multi-phase framework that trains a Skill Optimizer outside a frozen executor. The optimizer proposes skills, a frozen agent executes each repeatedly, and a programmatic verifier scores the outcomes. The scores provide relative credit and select mixed-outcome records. Matched successful and failed trajectories from these records form the next phase's refinement states, so the optimizer learns from the consequences of its earlier outputs. On BFCL v4 multi-turn and tau2-bench, WER improves average Pass@1 over the no-skill baseline by 7.80 and 3.85 points, respectively. Under an identical refinement workflow, it outperforms the same backbone without optimizer training by 9.35 and 10.29 points. The trained 4B optimizer reaches 76.63 percent on BFCL v4, outperforming all evaluated off-the-shelf general-purpose models used as skill optimizers on average.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑