arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MCP-Universe RL:一种通过强化学习训练MCP工具使用智能体的框架

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

Ziyang Luo, Yan Yang, Xiangru Jian, Ziji Shi, Xiaoqiang Lin, Jun Hao Liew, Silvio Savarese, Junnan Li

arXiv 2608.22167首次发表:更新:

发表机构

Salesforce AI Research(Salesforce AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出MCP-U RL开源框架,依托MCP协议解决RL训练中环境搭建与GPU利用率问题,在gpt-oss-20b上训练三类工具使用智能体并提升了任务奖励。

AI 中文摘要

强化学习(RL)已成为提升大语言模型(LLM)工具使用能力的有效方式,但多数现有RL框架仅停留在策略更新层面。对于每个新领域,用户需解决两个棘手的系统问题:一是为数百条并发轨迹搭建独立环境并将其与训练流程连接;二是安排 rollout 调度,确保GPU在多轮长对话(大量时间因工具调用缓慢而停滞)中保持忙碌。我们提出MCP-Universe RL(MCP-U RL),这是一个接管上述两项工作的开源框架。它采用模型上下文协议(MCP)作为与环境的接口,因此任何已作为MCP服务器暴露的工具无需特定于RL的集成代码即可接入训练。它构建了两个缺失的层并在各领域复用:环境编排层,通过可插拔容器后端配置、隔离和回收MCP环境;rollout编排层,其分阶段流水线可重叠轨迹,确保智能体在等待工具时GPU保持忙碌。后端无关的训练层则通过现有RL后端应用更新,已集成veRL和slime。仅修改任务规范的单一配置,我们在gpt-oss-20b上训练了软件工程、深度研究和通用工具使用智能体,且所有三类智能体的任务奖励均得到提升。

英文摘要

Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL frameworks stop at the policy update. For every new domain, the user is left with two hard systems problems: standing up an isolated environment for each of hundreds of concurrent trajectories and connecting it to training, and scheduling the rollout so that the GPU stays busy across long, multi-turn episodes that spend much of their time stalled on slow tool calls. We present MCP-Universe RL (MCP-U RL), an open-source framework that takes over both. It uses the Model Context Protocol (MCP) as the interface to the environment, so any tool already exposed as an MCP server plugs into training with no RL-specific integration code. It builds the two missing layers once and reuses them across domains: an environment-orchestration layer that provisions, isolates, and recycles the MCP environments over a pluggable container backend, and a rollout-orchestration layer whose staged pipeline overlaps trajectories to keep the GPU busy while episodes wait on tools. A backend-agnostic training layer then applies the update through an existing RL backend, with veRL and slime integrations. With one configuration, changing only the task specification, we train software-engineering, deep-research, and general tool-use agents on gpt-oss-20b and improve task reward in all three.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑