可控且可验证的工具使用数据合成用于代理强化学习
Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文提出COVERT框架,通过自进化合成和多级验证生成可靠的基础工具使用轨迹,并通过保留oracle工具调用和最终答案的增强方法提升环境复杂度,从而提升强化学习中工具调用策略的鲁棒性。
中文摘要 AI 辅助
现有合成工具使用数据集主要针对离线监督微调设计,而强化学习(RL)需要支持可验证在线回放的可执行环境。我们提出COVERT,一个两阶段流程:首先通过自进化合成和多级验证生成可靠的基线工具使用轨迹,然后应用保留oracle的增强方法,系统性地增加环境复杂性。这些增强方法引入干扰工具、间接或模糊的用户查询以及噪声、多格式或错误的工具输出,同时严格保留oracle工具调用和最终答案作为地面真实。此设计通过参考匹配自动计算奖励,对标准情况支持自动奖励计算,对特殊行为如错误检测支持轻量级裁判辅助验证,支持RL优化工具调用策略。在Qwen2.5-Instruct-14B上,COVERT-RL在BFCL v3上的整体准确率从56.5提升至59.9,在ACEBench上从53.0提升至59.3,对通用能力基准的回归最小;当叠加SFT时,进一步达到62.1和61.8,证实了加性收益。这些结果表明,保留oracle的合成环境为RL细化阶段提供了实用方法,补充SFT,以在模糊性和不可靠工具反馈下提升工具使用鲁棒性。
英文摘要
Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable environments that support reward-checkable online rollouts. We propose COVERT, a two-stage pipeline that first generates reliable base tool-use trajectories through self-evolving synthesis with multi-level validation, and then applies oracle-preserving augmentations that systematically increase environmental complexity. These augmentations introduce distractor tools, indirect or ambiguous user queries, and noisy, multi-format, or erroneous tool outputs, while strictly preserving oracle tool calls and final answers as ground truth. This design enables automatic reward computation via reference matching for standard cases and lightweight judge-assisted verification for special behaviors such as error detection, supporting RL optimization of tool-calling policies. On Qwen2.5-Instruct-14B, COVERT-RL improves overall accuracy on BFCL v3 from 56.5 to 59.9 and on ACEBench from 53.0 to 59.3, with minimal regressions on general-ability benchmarks; when stacked on SFT, it further reaches 62.1 and 61.8, confirming additive gains. These results suggest that oracle-preserving synthetic environments offer a practical RL refinement stage, complementary to SFT, for improving tool-use robustness under ambiguity and unreliable tool feedback.
发表机构
- The Pennsylvania State University(宾夕法尼亚州立大学)
- Amazon(亚马逊)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。