arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.09813cs.AI

可控且可验证的工具使用数据合成用于代理强化学习

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning

Siyuan Xu, Shiyang Li, Xin Liu, Tianyi Liu, Yixiao Li, Zhan Shi, Zixuan Zhang, Zilong Wang, Qingyu Yin, Jianshu Chen, Tuo Zhao, Bing Yin

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出COVERT框架,通过自进化合成和多级验证生成可靠的基础工具使用轨迹,并通过保留oracle工具调用和最终答案的增强方法提升环境复杂度,从而提升强化学习中工具调用策略的鲁棒性。

中文摘要 AI 辅助

现有合成工具使用数据集主要针对离线监督微调设计,而强化学习(RL)需要支持可验证在线回放的可执行环境。我们提出COVERT,一个两阶段流程:首先通过自进化合成和多级验证生成可靠的基线工具使用轨迹,然后应用保留oracle的增强方法,系统性地增加环境复杂性。这些增强方法引入干扰工具、间接或模糊的用户查询以及噪声、多格式或错误的工具输出,同时严格保留oracle工具调用和最终答案作为地面真实。此设计通过参考匹配自动计算奖励,对标准情况支持自动奖励计算,对特殊行为如错误检测支持轻量级裁判辅助验证,支持RL优化工具调用策略。在Qwen2.5-Instruct-14B上,COVERT-RL在BFCL v3上的整体准确率从56.5提升至59.9,在ACEBench上从53.0提升至59.3,对通用能力基准的回归最小;当叠加SFT时,进一步达到62.1和61.8,证实了加性收益。这些结果表明,保留oracle的合成环境为RL细化阶段提供了实用方法,补充SFT,以在模糊性和不可靠工具反馈下提升工具使用鲁棒性。

英文摘要

Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable environments that support reward-checkable online rollouts. We propose COVERT, a two-stage pipeline that first generates reliable base tool-use trajectories through self-evolving synthesis with multi-level validation, and then applies oracle-preserving augmentations that systematically increase environmental complexity. These augmentations introduce distractor tools, indirect or ambiguous user queries, and noisy, multi-format, or erroneous tool outputs, while strictly preserving oracle tool calls and final answers as ground truth. This design enables automatic reward computation via reference matching for standard cases and lightweight judge-assisted verification for special behaviors such as error detection, supporting RL optimization of tool-calling policies. On Qwen2.5-Instruct-14B, COVERT-RL improves overall accuracy on BFCL v3 from 56.5 to 59.9 and on ACEBench from 53.0 to 59.3, with minimal regressions on general-ability benchmarks; when stacked on SFT, it further reaches 62.1 and 61.8, confirming additive gains. These results suggest that oracle-preserving synthetic environments offer a practical RL refinement stage, complementary to SFT, for improving tool-use robustness under ambiguity and unreliable tool feedback.

发表机构

  • The Pennsylvania State University(宾夕法尼亚州立大学)
  • Amazon(亚马逊)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑