arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.12268cs.AI

CM2:基于检查清单奖励的多轮多步代理工具使用的强化学习

CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

  • University of California, Santa Barbara(加州大学圣巴巴拉分校)
  • Zoom Video Communications(Zoom视频通信公司)
  • University of Central Florida(中佛罗里达大学)
  • University of California, Los Angeles(加州大学洛杉矶分校)
  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

Zhen Zhang, Kaiqiang Song, Xun Wang, Yebowen Hu, Weixiang Yan, Chenyang Zhao, Henry Peng Zou, Haoyun Deng, Sathish Reddy Indurthi, Shujian Liu, Simin Ma, Xiaoya… 展开作者

Zhen Zhang, Kaiqiang Song, Xun Wang, Yebowen Hu, Weixiang Yan, Chenyang Zhao, Henry Peng Zou, Haoyun Deng, Sathish Reddy Indurthi, Shujian Liu, Simin Ma, Xiaoyang Wang, Xin Eric Wang, Song Wang

更新

中文总结 AI 辅助

CM2通过检查清单奖励提升多轮多步骤代理工具使用的强化学习效果,实现比监督微调更优的性能表现。

中文摘要 AI 辅助

人工智能代理越来越多地被用来通过推理多轮用户交互并调用外部工具来解决现实任务。然而,将强化学习应用于此类设置仍然具有挑战性:现实目标通常缺乏可验证的奖励,而是强调开放性行为;此外,针对多轮、多步骤代理工具使用的强化学习仍处于探索阶段;并且构建和维护可执行工具环境成本高昂,限制了规模和覆盖范围。我们提出了CM2,一种强化学习框架,它用检查清单奖励取代可验证的结果奖励。CM2将每个回合的预期行为分解为细粒度的二元标准,具有明确的证据基础和结构化元数据,将开放性判断转化为更稳定的分类式决策。为了在稳定性和信息性之间取得平衡,我们的方法采用稀疏奖励分配但密集评估标准的策略。训练是在可扩展的LLM模拟工具环境中进行的,避免了为大型工具集进行重工程。实验表明,CM2在监督微调上持续改进。从8B基础模型开始,训练于8k示例的强化学习数据集上,CM2在tau^-Bench上比SFT对照组提高8分,在BFCL-V4上提高10分,在ToolSandbox上提高12分。结果与或甚至优于同样规模的开源基线,包括判断模型。因此,CM2提供了一种可扩展的方法,用于优化多轮、多步骤工具使用代理,而无需依赖可验证的奖励。开源社区提供的代码:https://github.com/namezhenzhang/CM2-RLCR-Tool-Agent。

英文摘要

AI agents are increasingly used to solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. However, applying reinforcement learning to such settings remains difficult: realistic objectives often lack verifiable rewards and instead emphasize open-ended behaviors; moreover, RL for multi-turn, multi-step agentic tool use is still underexplored; and building and maintaining executable tool environments is costly, limiting scale and coverage. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, our method adopts a strategy of sparse reward assignment but dense evaluation criteria. Training is performed in a scalable LLM-simulated tool environment, avoiding heavy engineering for large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on tau^-Bench, by 10 points on BFCL-V4, and by 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus provides a scalable recipe for optimizing multi-turn, multi-step tool-using agents without relying on verifiable rewards. Code provided by the open-source community: https://github.com/namezhenzhang/CM2-RLCR-Tool-Agent.

↑