ToolVerse:为智能强化学习解锁大规模环境和长时任务
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究针对LLM智能体在复杂环境中工具集成问题,提出ToolVerse框架。通过构建大规模训练环境、设计任务策略及提出新算法,经实验验证该框架能增强LLM长时工具使用能力,提升性能与推理能力。
中文摘要 AI 辅助
虽然大语言模型(LLM)智能体在紧凑且定义明确的场景中展现出强大的推理能力,但面对需要无缝工具集成的大规模、多样且动态的现实世界环境时,它们难以保持稳健性和有效性。为填补这一差距,我们引入ToolVerse,这是一个全面的框架,可扩展智能强化学习环境,并使智能体在工具集成推理(TIR)任务中执行复杂的长时推理。首先,ToolVerse从近400个包含约4500个工具的现实世界模型上下文协议(MCP)自动构建大规模可执行智能体训练环境。其次,我们提出基于工具依赖图的任务设计策略,利用动态解锁采样算法生成长时任务,并生成GUST(图解锁采样任务)数据集。第三,为缓解长时智能强化学习中的信用分配问题,我们提出细粒度回合感知相对优势算法。我们使用ToolVerse进行了广泛的智能强化学习训练,并在多个智能基准上评估了我们的框架。实验结果表明,我们的框架显著增强了大语言模型在长时工具使用方面的能力,实现了显著的性能提升,并在动态环境中展示了强大的推理能力。
英文摘要
While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool integration. To address this gap, we introduce ToolVerse, a comprehensive framework that scales up agentic RL environments and enables agents to perform complex long-horizon reasoning in Tool-Integrated Reasoning (TIR) tasks. First, ToolVerse automatically builds the massive executable agent training environments from nearly 400 real-world Model Context Protocols (MCPs) that contain about 4500 tools. Second, we propose a task design strategy based on a tool dependency graph, utilizing Dynamic Unlocking Sampling Algorithm to generate long-horizon tasks, and produce GUST (Graph Unlocking Sampling Tasks) dataset. Third, to alleviate the credit assigment problem in long-horizon agentic RL, we propose a fine-grained Turn-Aware Relative Advantage algorithm. We conduct extensive Agentic RL training using ToolVerse and evaluate our framework on serveral agentic benchmarks. Experimental results demonstrate that our framework significantly strengthens LLMs' capabilities in long-horizon tool use, achieving a marked performance boost and showcasing robust reasoning within dynamic environments.
发表机构
- LongCat Interaction Team, Meituan(美团长猫互动团队)
- Peking University(北京大学)
- Fudan University(复旦大学)
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。