arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20082cs.LGcs.AIcs.CL

MATCH:具有课程调度和分层门控奖励的模型感知工具学习

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

发表机构中国科学院大学 · 荣耀终端有限公司
查看机构详情
  • University of the Chinese Academy of Sciences(中国科学院大学)
  • Honor Device Co., Ltd(荣耀终端有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Shihao Liu, Hao Yin, Lijun Liu, Zhengzong Chen, Yuanyuan Zhao, Fei Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对工具学习中固定阈值课程错位和奖励信用泄漏问题,提出MATCH框架,采用模型感知课程学习与分层门控奖励,在API-Bank和BFCL V3上取得领先准确率。

中文摘要 AI 辅助

工具学习使大型语言模型(LLM)能够使用外部工具完成超出参数化知识的任务。强化学习可以根据反馈优化工具调用行为,但当前方法仍面临两个问题:固定阈值的课程可能与策略不断演变的能力边界错位,以及当预测的工具错误时,加性奖励可能会泄漏参数级别的信用。为解决这些问题,我们提出了MATCH,一个用于模型感知工具学习的闭环框架,具有课程调度和分层门控奖励。模型感知课程学习(MACL)维护由奖励派生的样本难度,该难度与策略共同演变,每个epoch选择当前能力边界附近的样本以及一个包含较难案例的top-k池。分层工具调用门控奖励(HTGR)将工具名称、参数键和参数值作为门控链进行评分,仅在先决条件满足时在每个级别授予信用。相同的HTGR奖励驱动GRPO更新和MACL的难度刷新,从而在策略优化和样本调度之间形成闭环。在API-Bank和BFCL V3上,MATCH分别达到72.19%和62.87%的整体准确率,优于主要的监督和基于RL的基线。骨干实验进一步表明,在两个模型家族的四个骨干上均有一致的改进。

英文摘要

Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy's evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong. To address these problems, we propose MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards. Model-Aware Curriculum Learning (MACL) maintains reward-derived sample difficulty that co-evolves with the policy, and each epoch selects samples near the current capability boundary together with a top-k pool of harder cases. Hierarchical Tool-call Gated Reward (HTGR) scores tool name, argument key, and argument value as a gated chain, granting credit at each level only when prerequisites hold. The same HTGR rewards drive both GRPO updates and MACL's difficulty refresh, closing the loop between policy optimization and sample scheduling. On API-Bank and BFCL V3, MATCH reaches 72.19% and 62.87% overall accuracy, outperforming the main supervised and RL-based baselines. Backbone experiments further show consistent improvements across four backbones from two model families.

↑