arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30906cs.CL

ToolSearcher:通过强化学习在大规模工具选择中进行优化

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen

首次发表
浏览论文内容

中文总结 AI 辅助

ToolSearcher提出一种强化学习框架,通过类别约束区分、事件级搜索建模和轨迹对齐信用分配,解决大规模工具选择中的多轮搜索与细粒度优化问题,并在基准上超越强基线。

中文摘要 AI 辅助

大型语言模型(LLMs)在自然语言处理方面表现出色,但在与外部环境交互方面存在困难。工具学习为将LLMs扩展为可操作的智能体提供了一种有前景的方式,其中工具选择是成功使用工具的关键前提。现有工作通常假设工具集较小或预定义,导致大规模工具选择问题未被充分探索。现实世界的仓库包含大量且多样化的工具,使得LLMs在上下文长度限制下难以有效搜索、区分和组合工具。我们将大规模工具选择视为智能体强化学习的一个新挑战,并指出现有的基于知识问答的强化学习方法在考虑工具兼容性的同时选择工具方面存在不足。为解决这一挑战,我们提出了ToolSearcher,一种新颖的强化学习框架,用于在大规模工具选择中进行有效的多轮搜索和细粒度优化。具体来说,我们引入了类别约束的工具区分,以提高模型区分功能相似工具的能力;事件级搜索建模,以在多轮搜索中显式优化目标工具的发现;以及轨迹对齐的信用分配,为搜索-选择过程的不同阶段提供细粒度的奖励信号。在大规模工具选择基准上的大量实验表明,ToolSearcher在涉及迭代搜索和复杂工具组合的挑战性设置中始终优于一系列强基线。

英文摘要

Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.

发表机构

  • Zhejiang University(浙江大学)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑